Cerebras Systems
Principal SRE - AI Inference
Sunnyvale, CA · Principal
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- Architect self-service platforms and internal tooling that let product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
- 15+ years in SRE, infrastructure engineering, or platform engineering, with a record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.
- Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.
Responsibilities
- Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
- Mentor senior SREs, support critical incident escalations, and use production pain points to prioritize the highest-leverage automation work.
- Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Requirements
- Required Experience & Skills
- Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation.
- Hands-on experience with production observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, dashboards, and operational review loops.
- Experience with Bazel or other large-scale build systems in production.
Skills
- Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
Company info
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.