Cerebras Systems
Senior Software Development Engineer in Test (SDET) - AI Cluster
Toronto Office · Senior
Sponsorship not specifiedDetected 7 days ago
PythonGoDistributed SystemsAWSKubernetesPrometheusGrafanaMachine LearningData ScienceCompliance
About the role
- This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
- OpenAI recently announced a multi-year partnership https://openai.com/index/cerebras-partnership/ with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Responsibilities
- Automate first approach - In large scale deployment, automation drives efficiency and scalability. Aim for 100% automated tests to test all cluster features in areas of high availability, failure scenarios, performance, stress and security.
- In large scale deployment, automation drives efficiency and scalability. Aim for 100% automated tests to test all cluster features in areas of high availability, failure scenarios, performance, stress and security.
Requirements
- Bachelor's or master's degree in engineering in computer science, electrical, AI, data science or related field.
- 5+ years of experience in testing one of areas like enterprise software, distributed systems, datacenter hardware and software.
- Experience with debugging tools like pdb, gdb, strace and network monitors.
- Strong understanding of operating systems internals like memory management, file system working, security and performance.
- Strong understanding of datacenter layout, device performance characteristics like Servers, Memory, BIOS, PCIe, networking and storage.
- Strong coding skills in one of the programming languages like python, golang and C/C++.
- Strong debugging skills to debug issues in large distributed systems, hardware, and software. Experience with debugging tools like pdb, gdb, strace and network monitors.
- Experience with cloud technologies like AWS, kubernetes and dockers. Monitoring tools like grafana, prometheus is huge plus.
- Why Join Cerebras
- People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we've reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
Nice to have
- Experience with cloud technologies like AWS, kubernetes and dockers.
- Monitoring tools like grafana, prometheus is huge plus.
- Understanding and experience of ML model training and inference is a huge plus.
- Understand of ML hardware accelerators like GPU, custom accelerator ASIC is a huge plus.
Skills
- Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
Company info
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
- Cerebras is growing and innovating at a rapid pace and so is the ML community and AI models. Be a quick learner, adapt to new technologies, and bring your expertise. We are looking to hire a team with a diverse skill set.
Apply directly at Cerebras Systems →Create a free account for alerts like thisView Cerebras Systems immigration profile
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.