Cerebras Systems

Cerebras Systems

Staff Inference ML Runtime Engineer

Sunnyvale CA or Toronto Canada · Staff+

Sponsorship not specifiedDetected 97 days ago
PythonC++Machine LearningDeep LearningPyTorchLLMsComplianceCommunicationProblem Solving

About the role

  • Our mission is to empower enterprises, developers, and researchers to unlock the full potential of our platform, leveraging its performance, scalability, and flexibility.
  • As a Senior Software Engineer on the Inference ML Engineering team, you will play a key role in designing and implementing APIs, ML features, and tools that enable running state-of-the-art generative AI models on our custom hardware.
  • You will architect solutions that enable seamless model translation and execution, ensuring high throughput and low latency, while maintaining ease of use.

Responsibilities

  • Design and implement ML features (e.g., structured outputs, biased sampling, predicted outputs) that improve performance of generative AI models at inference time.
  • Design and implement high-throughput, low-latency multimodal inference models that support delivery of image, audio, and video inputs and outputs.
  • Maintain our scalable serving backend for handling many concurrent requests per minute.
  • Optimize software to accelerate generative LLM inference by achieving high throughput and low latency.
  • Evaluate trade-offs between different approaches, clearly articulate design choices, and develop detailed proposals for implementing new features.
  • Build and maintain robust automated test suites to ensure software quality, performance, and reliability.
  • Lead cross-functional initiative across the company to deliver high-quality inference solutions.

Requirements

  • Proficiency in Python for building and maintaining scalable systems.

Skills

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Mathematics, or a related field.
  • 8+ years of experience in large-scale software engineering, with a focus on deep learning or related domains.
  • Advanced proficiency in C++, with an emphasis on multi-threaded programming, performance optimization, and system-level development.
  • Demonstrated experience driving cross-functional projects.
  • Familiarity with LLM serving frameworks, such as vLLM, SGLang, and TensorRT-LLM.
  • Solid understanding of software architectural patterns for large-scale, high-performance applications.
  • Hands-on experience with ML frameworks, such as PyTorch, and a strong understanding of their underlying architectures.
  • Strong problem-solving skills, with the ability to balance technical depth with practical implementation constraints.

Benefits

  • Drive and provide technical guidance to a team of software engineers working on complex machine learning integration projects.
  • Stay up-to-date with advancements in machine learning and deep learning, and apply state-of-the-art techniques to enhance our solutions.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.