Cerebras Systems
Staff Inference ML Runtime Engineer
Sunnyvale CA or Toronto Canada · Staff+
Sponsorship not specifiedDetected 97 days ago
PythonC++Machine LearningDeep LearningPyTorchLLMsComplianceCommunicationProblem Solving
About the role
- Our mission is to empower enterprises, developers, and researchers to unlock the full potential of our platform, leveraging its performance, scalability, and flexibility.
- As a Senior Software Engineer on the Inference ML Engineering team, you will play a key role in designing and implementing APIs, ML features, and tools that enable running state-of-the-art generative AI models on our custom hardware.
- You will architect solutions that enable seamless model translation and execution, ensuring high throughput and low latency, while maintaining ease of use.
Responsibilities
- Design and implement ML features (e.g., structured outputs, biased sampling, predicted outputs) that improve performance of generative AI models at inference time.
- Design and implement high-throughput, low-latency multimodal inference models that support delivery of image, audio, and video inputs and outputs.
- Maintain our scalable serving backend for handling many concurrent requests per minute.
- Optimize software to accelerate generative LLM inference by achieving high throughput and low latency.
- Evaluate trade-offs between different approaches, clearly articulate design choices, and develop detailed proposals for implementing new features.
- Build and maintain robust automated test suites to ensure software quality, performance, and reliability.
- Lead cross-functional initiative across the company to deliver high-quality inference solutions.
Requirements
- Proficiency in Python for building and maintaining scalable systems.
Skills
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Mathematics, or a related field.
- 8+ years of experience in large-scale software engineering, with a focus on deep learning or related domains.
- Advanced proficiency in C++, with an emphasis on multi-threaded programming, performance optimization, and system-level development.
- Demonstrated experience driving cross-functional projects.
- Familiarity with LLM serving frameworks, such as vLLM, SGLang, and TensorRT-LLM.
- Solid understanding of software architectural patterns for large-scale, high-performance applications.
- Hands-on experience with ML frameworks, such as PyTorch, and a strong understanding of their underlying architectures.
- Strong problem-solving skills, with the ability to balance technical depth with practical implementation constraints.
Benefits
- Drive and provide technical guidance to a team of software engineers working on complex machine learning integration projects.
- Stay up-to-date with advancements in machine learning and deep learning, and apply state-of-the-art techniques to enhance our solutions.
Company info
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
Apply directly at Cerebras Systems →Create a free account for alerts like thisView Cerebras Systems immigration profile
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.