Zyphra
Platform Engineer
San Francisco
Sponsorship not specifiedDetected 127 days ago
Backend DevelopmentAWSGCPCloud PlatformsDockerKubernetesTerraformAnsibleCI/CDDevOpsPlatform EngineeringMachine LearningIncident ResponseLoad TestingResearchExperimental Design
About the role
- As a Platform Engineer, you'll be responsible for designing and maintaining the systems that keep Zyphra's infrastructure robust, observable, secure, and scalable.
- Your work will be essential to ensuring the reliability and reproducibility of ML workloads, the safety and control of deployments, and the long-term maintainability of our compute environments.
Responsibilities
- Designing resilient build and deployment systems across research and production environments
- Implementing secure release processes with strong auditability and rollback support
- Relocation and immigration support on a case-by-case basis
Requirements
- Experience in high-performance compute environments, such as ML clusters or GPU farms as well as hyperscaler cloud environments (i.e. AWS, GCP, etc.)
- Familiarity with containers (i.e., Docker, Apptainer) and their integration with scheduling systems (i.e., Kubernetes, Slurm)
- Experience managing run-books, DRP, change management, and general fault tolerance
- Experience with deployment strategies at scale
- Experience designing reliable environments for experimental workloads and reproducible runs
- Knowledge of compliance and audit standards in deployment and system security
- Experience with load testing, fault injection, and chaos engineering to harden systems under stress
- Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued
Nice to have
- Familiarity with software release engineering for ML/AI systems is a plus
Skills
- Experience with infrastructure as code (e.g., Ansible, Terraform)
- Prior work supporting ML/AI infrastructure, including GPU management and workload optimization
- Exposure to backend development for ML model serving (i.e., vLLM, Ray, SGLang, Triton)
- We strongly value new and crazy ideas and are very willing to bet big on new ideas
- We move as quickly as we can; we aim to minimize the bar to impact as low as possible
- We all enjoy what we do and love discussing AI
- Building and improving observability systems (monitoring, logging, alerting)
- Managing Infrastructure as a Service across the stack along with CI/CD in close partnership with engineering teams
- Collaborating closely with ML engineers, DevOps, and infra teams to improve system reliability and performance
- Leading incident response, root-cause analysis, and postmortems with a focus on learning and prevention
Compensation
- Competitive compensation and 401(k) plan
Benefits
- Comprehensive medical, dental, vision, and FSA plans
- Unlimited PTO and company holidays
Company info
- ZYPHRA IS AN ARTIFICIAL INTELLIGENCE COMPANY BASED IN SAN FRANCISCO, CALIFORNIA.
This listing is sourced directly from Zyphra's careers page and normalized into a canonical job model.