Gradial

Gradial

Principal Platform Engineer

Seattle, WA · Principal

Sponsorship not specifiedDetected 6 days ago
TypeScriptPythonCloud PlatformsKubernetesCI/CDDevOpsSite Reliability EngineeringPlatform EngineeringMachine LearningIncident ResponseLeadershipProblem Solving

About the role

  • As a Principal Platform Engineer at Gradial, you will shape the foundation our platform runs on as we scale.
  • You will work closely with the CTO and engineering team to make our systems faster, more resilient, and easier to operate in a high-growth environment.
  • This is a hands-on individual contributor leadership role for someone who wants real ownership, high leverage, and the opportunity to define how platform reliability looks at an AI-native company.

Responsibilities

  • Set the standard for how we design, ship, and operate reliable systems.
  • Build the tooling and automation that help engineers move faster with more confidence.
  • Drive improvements in monitoring, alerting, incident response, and service readiness.
  • Partner with engineering to identify scaling risks early and solve them before they slow us down.
  • A front-row seat to building category-defining AI infrastructure
  • Have an innate drive for being 1% better than the day before, building towards greatness.
  • Thrive in fast-paced, hyper-growth environments where building better > maintaining status quo.

Requirements

  • 5+ years of experience in platform engineering, infrastructure, SRE, DevOps, or related roles with direct ownership of production systems.
  • Proven success designing and operating production-grade infrastructure in fast-moving, high-growth environments.
  • Deep expertise in Kubernetes, cloud-native architecture, and container orchestration.
  • Strong experience with infrastructure as code, GitOps, CI/CD workflows, and modern deployment practices.
  • A track record of leading through influence, making sound technical decisions, and raising the bar across engineering teams.

Nice to have

  • Familiarity with AI or ML infrastructure, including GPU provisioning, model deployment, or compute-intensive workloads.
  • Experience supporting cloud or multi-cloud environments with a focus on resilience and scale.
  • Comfort with TypeScript or Python for internal tooling and operational automation.
  • Fast-paced environment with flexibility and ownership
  • Real impact, zero bureaucracy
  • Learn quickly, actively seek out new challenges, and regularly reconsider "how it's always been done."
  • Embrace AI as a core tool for problem-solving, innovation, and scale.
  • Show customer-obsession (internal or external), high ownership/accountability, and bias for action.

Compensation

  • Competitive salary and meaningful equity

Benefits

  • Competitive salary and meaningful equity
  • Comprehensive health, dental and vision coverage

This listing is sourced directly from Gradial's careers page and normalized into a canonical job model.