Cerebras Systems

Cerebras Systems

Compute Server Platform Architect

Sunnyvale CA or Toronto Canada

Sponsorship not specifiedDetected 97 days ago
PythonC++LinuxMachine LearningLLMsComplianceBudgetingLeadershipCommunication

About the role

  • As a Compute / Server Platform Architect on the Cluster Architecture Team, you will own the server-side platform architecture that enables Cerebras CS3-based AI clusters (training and inference) to deliver predictable performance, scalability, and reliability.
  • Our accelerators are network-attached, so the x86 server fleet is a first-class part of the end-to-end system: it runs critical-path runtime functions (for example orchestration, prompt caching, and IO/control services) and must be co-designed with software for token-level latency, throughput, and cost efficiency.
  • You will translate workload behavior into CPU, memory, IO, PCIe, and host-networking requirements, drive platform evaluations with vendors, and provide technical leadership through qualification and production adoption in close partnership with other function leaders and TPMs.

Responsibilities

  • Own the architecture for all server roles in Cerebras clusters, including definitions of server types, configurations, and lifecycle strategy.
  • Define and maintain server formulas (counts and ratios per CS-3 count, cluster size, and workload type) including capacity planning and headroom policy.
  • Develop performance and scaling models
  • validate with microbenchmarks and workload-level experiments
  • identify bottlenecks and drive cross-stack fixes.
  • Lead technical vendor engagements (OEM/ODM and component vendors): influence roadmap, request platform knobs, and drive joint debugging on performance or reliability issues.
  • Define qualification and acceptance criteria (performance, stability, operability) and partner with the Infrastructure Hardware TPM to execute qualification plans and land changes cleanly into production.
  • Support bring-up and rare deployment debugging in lab and staging environments
  • drive root-cause analysis for regressions spanning firmware, drivers, OS, and runtime behavior.
  • Develop performance and scaling models; validate with microbenchmarks and workload-level experiments; identify bottlenecks and drive cross-stack fixes.

Requirements

  • PhD. in Computer Science or Electrical/Computer Engineering and + 8 years industry experience, or Master's/Bachelor's in CS or EE + 10 years industry experience.
  • 5+ years of experience in server platform architecture, systems performance engineering, or large-scale infrastructure design for AI/ML, HPC, or performance-sensitive distributed systems.
  • Experience reasoning about high-performance IO paths, including NIC behavior at a systems level, RDMA/RoCE concepts, and NVMe performance characteristics.
  • Familiarity with application and system software (C, C++, Python).

Skills

  • Five Reasons to Join Cerebras in 2026.
  • Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.
  • This website or its third-party tools process personal data.
  • For more details, click here to review our CCPA disclosure notice.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
  • At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.