Cerebras Systems
Sr./Staff TPM - Inference Capacity
Headquarters/Sunnyvale Office · Staff+
Sponsorship not specifiedDetected 22 days ago
PythonSQLCloud PlatformsGrafanaSite Reliability EngineeringMachine LearningComplianceJiraConfluenceSalesPipeline ManagementForecastingProcess ImprovementLeadership
About the role
- As demand for AI continues to accelerate, intelligent capacity management becomes one of the company's most strategic challenges.
- Every customer commitment, model launch, and infrastructure investment depends on making the right capacity decisions at the right time.
- This is a highly visible role working directly with Engineering, Product, Infrastructure, SRE, Operations, and executive leadership to maximize utilization of one of the world's most advanced AI inference fleets.
Responsibilities
- Run weekly capacity planning and daily capacity and deployment tracking with Engineering, product and operations team. Own fleet utilization reporting and forecasting
- Drive capacity planning for new customer deployments and major model launches
- Drive continuous improvement and stakeholder adoption of new capacity management platform
- Drive org level strategic initiatives related to capacity expansion, improving fleet efficiency and maximizing effective utilization of available systems
- Lead planning around major infrastructure events including but not limited to new customer commits, new model releases, change to DC/cluster architecture, etc. that impacts capacity and fleet utilization. Update capacity plans and forecasts accordingly.
- Maintain Jira EPICs and Confluence pages related to capacity planning, reporting and change management to ensure execution transparency across teams
- Strong data fluency: SQL, Grafana, basic Python or Flux to pull your own numbers without waiting for an analyst
Requirements
- 5+ years of TPM, technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning
- Experience leading large cross-functional programs involving Engineering, Product, and Operations
- Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling
- Track record of running a recurring cross-functional ritual involving senior engineers and LT
- Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium
Skills
- Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
- which customer tenants land where, which clusters absorb new launches, which freezes are in effect etc.
- Run the weekly capacity and utilization report for the Inference Service leadership.
- Drive capacity planning tool adoption.
- In general, Contribute to the continuous process improvement and development of internal capacity management tools.
- Incident tracking and postmortems.
- Proactively identify and mitigate capacity bottlenecks, risks, and dependencies.
Company info
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
- Partner closely with the SRE and product team to run the weekly capacity review across different customers/models/clusters.
Apply directly at Cerebras Systems →Create a free account for alerts like thisView Cerebras Systems immigration profile
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.