Cerebras Systems
Director/Sr. Manager, AI Inference Model Scaling
Sunnyvale, CA · Director
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- You will define the technical vision, organizational strategy, and execution roadmap for a globally distributed engineering team responsible for enabling the latest foundation models on Cerebras hardware.
- This role combines deep technical leadership with organizational excellence.
- The Inference Model Scaling team enables state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras' Wafer-Scale Engine (WSE).
Responsibilities
- Lead design reviews and establish engineering standards.
- Drive support for emerging LLM architectures and inference workloads.
- Develop future technical leaders and managers.
- Drive organizational planning, headcount strategy, and investment priorities.
- Partner with Cloud Platform, ML, and Hardware teams in planning and delivering for end-to-end service enablement in Cloud and On-Premise settings
Requirements
- 5+ years leading engineering teams.
- Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
- Strong understanding of graph compilation and optimization.
- Experience with Python and C++.
- Experience delivering production-quality software.
Nice to have
- Experience supporting PyTorch, JAX, TensorFlow, or ONNX.
- Experience with LLM inference or training systems.
- Familiarity with distributed compilation.
- Experience working with hardware architects.
- Experience leading teams through rapid growth.
- Why Join Cerebras
- With dozens of model releases and rapid growth, we've reached an inflection point in our business.
- Members of our team tell us there are five main reasons they joined Cerebras:
Skills
- Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
- Sunnyvale, CA or Toronto, Canada (Hybrid)
Company info
- We build the compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernel enablement, and runtime integration that together make next-generation AI models execute with industry-leading performance.
- The team works at the intersection of machine learning frameworks, compiler technologies, distributed systems, hardware architecture, and model optimization.
- We collaborate closely with hardware architects, runtime engineers, cloud platform teams, AI researchers, and strategic customers to rapidly bring new model architectures into production.
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.