Cerebras Systems
Cluster Operations Software Engineer
Sunnyvale, CA · Senior
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.
- You will play a critical role in ensuring the health, performance, and availability of our infrastructure, maximizing compute capacity, and supporting our growing AI initiatives.
- The ideal candidate is a proactive problem-solver with expertise in large-scale compute infrastructure, dependable and an advocate for customer success.
Responsibilities
- Build and own software solutions that power cluster operations, including monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.
- Collaborate with cross-functional teams to translate operational requirements into scalable O&M products and platform capabilities.
- Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management across global AI infrastructure.
- Manage and operate multiple advanced AI compute infrastructure clusters.
- Provide 24/7 monitoring and support, leveraging automated tools and performing hands-on troubleshooting as needed.
- Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
- Proficient in Python and Go, with experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments.
Requirements
- Experience and Expertise in distributed systems is a must.
- Extensive knowledge of Docker containers and container orchestration platforms like k8s.
- Proven ability to troubleshoot and resolve complex technical issues in a timely and efficient manner.
- Experience with monitoring and alerting systems.
- Ability to work effectively in a fast-paced environment.
- Knowledge of technologies like Ethernet, RoCE, TCP/IP, etc. is desired.
- Knowledge of cloud computing platforms (e.g., AWS, GCP, Azure).
Skills
- Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
Benefits
- Monitor and oversee cluster health, proactively identifying and resolving potential issues.
- 6-8 years of relevant experience in managing and operating complex compute infrastructure, preferably in the context of machine learning or high-performance computing.
Company info
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.