Zyphra

Zyphra

Platform Engineer

San Francisco

Sponsorship not specifiedDetected 127 days ago
Backend DevelopmentAWSGCPCloud PlatformsDockerKubernetesTerraformAnsibleCI/CDDevOpsPlatform EngineeringMachine LearningIncident ResponseLoad TestingResearchExperimental Design

About the role

  • As a Platform Engineer, you'll be responsible for designing and maintaining the systems that keep Zyphra's infrastructure robust, observable, secure, and scalable.
  • Your work will be essential to ensuring the reliability and reproducibility of ML workloads, the safety and control of deployments, and the long-term maintainability of our compute environments.

Responsibilities

  • Designing resilient build and deployment systems across research and production environments
  • Implementing secure release processes with strong auditability and rollback support
  • Relocation and immigration support on a case-by-case basis

Requirements

  • Experience in high-performance compute environments, such as ML clusters or GPU farms as well as hyperscaler cloud environments (i.e. AWS, GCP, etc.)
  • Familiarity with containers (i.e., Docker, Apptainer) and their integration with scheduling systems (i.e., Kubernetes, Slurm)
  • Experience managing run-books, DRP, change management, and general fault tolerance
  • Experience with deployment strategies at scale
  • Experience designing reliable environments for experimental workloads and reproducible runs
  • Knowledge of compliance and audit standards in deployment and system security
  • Experience with load testing, fault injection, and chaos engineering to harden systems under stress
  • Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued

Nice to have

  • Familiarity with software release engineering for ML/AI systems is a plus

Skills

  • Experience with infrastructure as code (e.g., Ansible, Terraform)
  • Prior work supporting ML/AI infrastructure, including GPU management and workload optimization
  • Exposure to backend development for ML model serving (i.e., vLLM, Ray, SGLang, Triton)
  • We strongly value new and crazy ideas and are very willing to bet big on new ideas
  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible
  • We all enjoy what we do and love discussing AI
  • Building and improving observability systems (monitoring, logging, alerting)
  • Managing Infrastructure as a Service across the stack along with CI/CD in close partnership with engineering teams
  • Collaborating closely with ML engineers, DevOps, and infra teams to improve system reliability and performance
  • Leading incident response, root-cause analysis, and postmortems with a focus on learning and prevention

Compensation

  • Competitive compensation and 401(k) plan

Benefits

  • Comprehensive medical, dental, vision, and FSA plans
  • Unlimited PTO and company holidays

Company info

  • ZYPHRA IS AN ARTIFICIAL INTELLIGENCE COMPANY BASED IN SAN FRANCISCO, CALIFORNIA.

This listing is sourced directly from Zyphra's careers page and normalized into a canonical job model.