Cerebras Systems

Cerebras Systems

Infrastructure Hardware Technical Program Manager (Server and Network Systems)

Sunnyvale CA or Toronto Canada

Sponsorship not specifiedDetected 97 days ago
LinuxMachine LearningLLMsComplianceLeadership

About the role

  • Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device.
  • OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
  • Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services.

Responsibilities

  • Own end-to-end program execution for server systems and network equipment in Cerebras clusters, including new platforms, refreshes, and major component/config changes.
  • Drive requirements gathering and convert inputs into executable plans with clear milestones, readiness gates, and cross-functional deliverables.
  • Build and manage integrated schedules across vendors and internal teams, track dependencies, critical path, and risks.
  • Manage OEM/ODM and switch/vendor engagements (RFI/RFP, samples, escalations, roadmap alignment).
  • Partner with Compute / Server Platform / Network Architects to turn architectural decisions into qualification plans, acceptance criteria, and rollout strategies.
  • Lead qualification and release readiness (lab/staging validation, regression tracking, go/no-go decisions).
  • Own risk and change management into production, including versioning, rollout sequencing, and stakeholder communication.
  • Ensure operational readiness with deployment and fleet teams and maintain alignment with rack/physical DC owners on power, cooling, space, and cabling constraints.
  • People who are serious about software make their own hardware.
  • Build a breakthrough AI platform beyond the constraints of the GPU.

Requirements

  • This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Skills

  • B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or equivalent experience.
  • Experience coordinating complex server and/or datacenter network programs across OEM/ODMs, switch vendors, and internal engineering teams.
  • Familiarity with Linux server fleet management (provisioning, firmware/BIOS, drivers, field triage).
  • Ability to operate in ambiguity and keep parallel server and network workstreams aligned.
  • Experience with AI/ML, HPC, or performance-sensitive distributed infrastructure is a plus.
  • integrated plans, risk management, dependency tracking, and executive-level communication.
  • Why Join Cerebras
  • Members of our team tell us there are five main reasons they joined Cerebras:

Company info

  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
  • At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry.
  • Apply today and become part of the forefront of groundbreaking advancements in AI!

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.