Cerebras Systems
Infrastructure Hardware Technical Program Manager (Server and Network Systems)
Sunnyvale CA or Toronto Canada
Sponsorship not specifiedDetected 97 days ago
LinuxMachine LearningLLMsComplianceLeadership
About the role
- Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device.
- OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
- Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services.
Responsibilities
- Own end-to-end program execution for server systems and network equipment in Cerebras clusters, including new platforms, refreshes, and major component/config changes.
- Drive requirements gathering and convert inputs into executable plans with clear milestones, readiness gates, and cross-functional deliverables.
- Build and manage integrated schedules across vendors and internal teams, track dependencies, critical path, and risks.
- Manage OEM/ODM and switch/vendor engagements (RFI/RFP, samples, escalations, roadmap alignment).
- Partner with Compute / Server Platform / Network Architects to turn architectural decisions into qualification plans, acceptance criteria, and rollout strategies.
- Lead qualification and release readiness (lab/staging validation, regression tracking, go/no-go decisions).
- Own risk and change management into production, including versioning, rollout sequencing, and stakeholder communication.
- Ensure operational readiness with deployment and fleet teams and maintain alignment with rack/physical DC owners on power, cooling, space, and cabling constraints.
- People who are serious about software make their own hardware.
- Build a breakthrough AI platform beyond the constraints of the GPU.
Requirements
- This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Skills
- B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or equivalent experience.
- Experience coordinating complex server and/or datacenter network programs across OEM/ODMs, switch vendors, and internal engineering teams.
- Familiarity with Linux server fleet management (provisioning, firmware/BIOS, drivers, field triage).
- Ability to operate in ambiguity and keep parallel server and network workstreams aligned.
- Experience with AI/ML, HPC, or performance-sensitive distributed infrastructure is a plus.
- integrated plans, risk management, dependency tracking, and executive-level communication.
- Why Join Cerebras
- Members of our team tell us there are five main reasons they joined Cerebras:
Company info
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
- At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry.
- Apply today and become part of the forefront of groundbreaking advancements in AI!
Apply directly at Cerebras Systems →Create a free account for alerts like thisView Cerebras Systems immigration profile
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.