Cerebras Systems

Cerebras Systems

Principal Engineer, AI Inference Reliability

Remote, California, United States; Sunnyvale CA or Toronto Canada · Principal

Sponsorship not specifiedDetected 97 days ago
PythonGoC++Distributed SystemsMachine LearningLLMsIncident ResponseComplianceLeadershipCommunication

About the role

  • Join us to help scale inference and accelerate AI.

Responsibilities

  • Define and drive reliability strategy: establish SLOs and ensure alignment across engineering.
  • Design and implement reliability mechanisms: build and evolve systems for fault detection, graceful degradation, failover, throttling, and recovery across multiple regions and data centers.
  • Lead large-scale incident management: own postmortems, root-cause analysis, and prevention loops for reliability-related incidents.
  • Architect for reliability and observability: influence system design for redundancy, durability, and debuggability.
  • Develop reliability tooling: create internal tools and frameworks for chaos testing, load simulation, and distributed fault injection.
  • Collaborate broadly: work across software, infrastructure, and hardware teams to ensure reliability is embedded into every layer of our inference service.
  • Deep and hard-earned experience of reliability principles: SLO/SLI/SLA design, incident response, and postmortem culture.
  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.
  • prior experience building large-scale AI infrastructure systems.
  • People who are serious about software make their own hardware.

Requirements

  • Bachelor's or master's degree in computer science or related field.
  • 7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
  • Strong programming skills in at least one popular backend programming language such as Python, C++, Go, or Rust.

Nice to have

  • With dozens of model releases and rapid growth, we've reached an inflection point in our business.

Skills

  • Excellent communication and cross-functional leadership skills.
  • Five Reasons to Join Cerebras in 2026.
  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • This website or its third-party tools process personal data.
  • For more details, click here to review our CCPA disclosure notice.

Compensation

  • In October 2025, we announced our series G funding, raising $1.1 billion USD to accelerate the expansion of our products and services to meet global AI demand.

Benefits

  • build dashboards and alerts that measure service health and provide actionable insights.

Company info

  • The Cerebras Inference team's mission is to deliver the world's most performant, secure, and reliable enterprise-grade AI service.
  • We build and operate large-scale distributed systems that power AI inference at unprecedented speed and efficiency.
  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
  • At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.