Cerebras Systems
Principal Engineer, AI Inference Reliability
Remote, California, United States; Sunnyvale CA or Toronto Canada · Principal
Sponsorship not specifiedDetected 97 days ago
PythonGoC++Distributed SystemsMachine LearningLLMsIncident ResponseComplianceLeadershipCommunication
About the role
- Join us to help scale inference and accelerate AI.
Responsibilities
- Define and drive reliability strategy: establish SLOs and ensure alignment across engineering.
- Design and implement reliability mechanisms: build and evolve systems for fault detection, graceful degradation, failover, throttling, and recovery across multiple regions and data centers.
- Lead large-scale incident management: own postmortems, root-cause analysis, and prevention loops for reliability-related incidents.
- Architect for reliability and observability: influence system design for redundancy, durability, and debuggability.
- Develop reliability tooling: create internal tools and frameworks for chaos testing, load simulation, and distributed fault injection.
- Collaborate broadly: work across software, infrastructure, and hardware teams to ensure reliability is embedded into every layer of our inference service.
- Deep and hard-earned experience of reliability principles: SLO/SLI/SLA design, incident response, and postmortem culture.
- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.
- prior experience building large-scale AI infrastructure systems.
- People who are serious about software make their own hardware.
Requirements
- Bachelor's or master's degree in computer science or related field.
- 7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
- Strong programming skills in at least one popular backend programming language such as Python, C++, Go, or Rust.
Nice to have
- With dozens of model releases and rapid growth, we've reached an inflection point in our business.
Skills
- Excellent communication and cross-functional leadership skills.
- Five Reasons to Join Cerebras in 2026.
- Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
- This website or its third-party tools process personal data.
- For more details, click here to review our CCPA disclosure notice.
Compensation
- In October 2025, we announced our series G funding, raising $1.1 billion USD to accelerate the expansion of our products and services to meet global AI demand.
Benefits
- build dashboards and alerts that measure service health and provide actionable insights.
Company info
- The Cerebras Inference team's mission is to deliver the world's most performant, secure, and reliable enterprise-grade AI service.
- We build and operate large-scale distributed systems that power AI inference at unprecedented speed and efficiency.
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
- At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry.
Apply directly at Cerebras Systems →Create a free account for alerts like thisView Cerebras Systems immigration profile
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.