Cerebras Systems

Cerebras Systems

Sr./Staff TPM - Inference Capacity

Headquarters/Sunnyvale Office · Staff+

Sponsorship not specifiedDetected 22 days ago
PythonSQLCloud PlatformsGrafanaSite Reliability EngineeringMachine LearningComplianceJiraConfluenceSalesPipeline ManagementForecastingProcess ImprovementLeadership

About the role

  • As demand for AI continues to accelerate, intelligent capacity management becomes one of the company's most strategic challenges.
  • Every customer commitment, model launch, and infrastructure investment depends on making the right capacity decisions at the right time.
  • This is a highly visible role working directly with Engineering, Product, Infrastructure, SRE, Operations, and executive leadership to maximize utilization of one of the world's most advanced AI inference fleets.

Responsibilities

  • Run weekly capacity planning and daily capacity and deployment tracking with Engineering, product and operations team. Own fleet utilization reporting and forecasting
  • Drive capacity planning for new customer deployments and major model launches
  • Drive continuous improvement and stakeholder adoption of new capacity management platform
  • Drive org level strategic initiatives related to capacity expansion, improving fleet efficiency and maximizing effective utilization of available systems
  • Lead planning around major infrastructure events including but not limited to new customer commits, new model releases, change to DC/cluster architecture, etc. that impacts capacity and fleet utilization. Update capacity plans and forecasts accordingly.
  • Maintain Jira EPICs and Confluence pages related to capacity planning, reporting and change management to ensure execution transparency across teams
  • Strong data fluency: SQL, Grafana, basic Python or Flux to pull your own numbers without waiting for an analyst

Requirements

  • 5+ years of TPM, technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning
  • Experience leading large cross-functional programs involving Engineering, Product, and Operations
  • Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling
  • Track record of running a recurring cross-functional ritual involving senior engineers and LT
  • Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium

Skills

  • Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
  • which customer tenants land where, which clusters absorb new launches, which freezes are in effect etc.
  • Run the weekly capacity and utilization report for the Inference Service leadership.
  • Drive capacity planning tool adoption.
  • In general, Contribute to the continuous process improvement and development of internal capacity management tools.
  • Incident tracking and postmortems.
  • Proactively identify and mitigate capacity bottlenecks, risks, and dependencies.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
  • Partner closely with the SRE and product team to run the weekly capacity review across different customers/models/clusters.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.