Cerebras Systems

Cerebras Systems

AI Infrastructure Operations Engineer

Sunnyvale CA or Toronto Canada

Sponsorship not specifiedDetected 97 days ago
PythonAWSGCPAzureDockerKubernetesLinuxMachine LearningLLMsComplianceTCP/IPCommunicationCollaboration

About the role

  • These clusters would provide the candidate an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.
  • You will play a critical role in ensuring the health, performance, and availability of our infrastructure, maximizing compute capacity, and supporting our growing AI initiatives.
  • The ideal candidate is a proactive problem-solver with expertise in large-scale compute infrastructure, dependable and an advocate for customer success.

Responsibilities

  • Manage and operate multiple advanced AI compute infrastructure clusters.
  • Provide 24/7 monitoring and support, leveraging automated tools and performing hands-on troubleshooting as needed.
  • Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
  • Contribute to the development and improvement of our monitoring and support processes.

Requirements

  • Experience with cross-functional team projects.
  • This role requires a deep understanding of Linux-based systems, containerization technologies, and experience with monitoring and troubleshooting complex distributed systems.

Skills

  • Strong proficiency in Python scripting for automation and system administration.
  • Deep understanding of Linux-based compute systems and command-line tools.
  • Extensive knowledge of Docker containers and container orchestration platforms like k8s and SLURM.
  • Proven ability to troubleshoot and resolve complex technical issues in a timely and efficient manner.
  • Experience with monitoring and alerting systems.
  • Excellent communication and collaboration skills.
  • Ability to work effectively in a fast-paced environment.
  • Willingness to participate in a 24/7 on-call rotation.
  • Operating large scale GPU clusters.
  • Knowledge of technologies like Ethernet, RoCE, TCP/IP, etc. is desired.
  • Knowledge of cloud computing platforms (e.g., AWS, GCP, Azure).
  • Familiarity with machine learning frameworks and tools.

Benefits

  • Monitor and oversee cluster health, proactively identifying and resolving potential issues.
  • 6-8 years of relevant experience in managing and operating complex compute infrastructure, preferably in the context of machine learning or high-performance computing.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.