Cerebras Systems

Cerebras Systems

Site Reliability Engineer - Ops & Automation

Sunnyvale CA or Toronto Canada · Staff+

Sponsorship not specifiedDetected 97 days ago
PythonGoKubernetesCI/CDPrometheusGrafanaSite Reliability EngineeringMachine LearningLLMsComplianceCollaborationMentoringArgo CD

About the role

  • This role offers immediate ownership of real production systems at a growing scale, direct mentorship from seasoned engineers, and close collaboration with incoming Staff SREs who will focus on long-term automation.
  • After ~1 month of shared hands-on operations with the Staff engineers, you'll primarily operate the current setup, bring up new capacity in high-stakes environments and help bring new continuous delivery pipelines into production use.
  • If you thrive in high-ownership SRE roles at scale and want to help shape a team from the ground up in cutting-edge AI Inference infrastructure, this is your chance.

Responsibilities

  • Remain hands-on with operational execution (releases, capacity changes, cluster upgrades) over the next year as we build robust continuous delivery pipelines and self-service capabilities
  • Build reusable automation and internal developer tools that minimize operational toil and cross-team friction
  • Develop and extend telemetry, observability and alerting solutions to ensure operational reliability at scale
  • Collaborate with Cluster Ops and development teams to identify high-impact automation opportunities and iterate quickly
  • Solid Python or Go for building tools and automation

Requirements

  • Required Experience & Skills

Nice to have

  • Contribute to reliability practices (SLOs, post-mortems, capacity planning)
  • 2-4+ years in SRE with a strong operations or automation focus
  • Production Kubernetes experience
  • Proficiency with Prometheus, Grafana, and observability-driven workflows
  • Ability to measure and communicate impact - reliability metrics, operational toil, velocity gains
  • Hands-on GitOps expertise, Argo CD / Flux or equivalent, is a plus
  • Experience with building continuous delivery pipelines is a strong plus
  • Experience with Bazel or similar build systems is a strong plus.

Skills

  • Five Reasons to Join Cerebras in 2026.
  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • This website or its third-party tools process personal data.
  • For more details, click here to review our CCPA disclosure notice.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.