xAI
ML Infrastructure Engineer
Palo Alto, California, United States
Sponsorship not specified$180k-$440kDetected 12 hours ago
PythonRustC++Distributed SystemsFull-Stack DevelopmentAnsibleLinuxMachine LearningDeep LearningPyTorchData EngineeringA/B TestingLeadershipCommunicationCollaborationMentoring
About the role
- We're looking for exceptional engineers who are passionate about our mission and have a strong desire to make a meaningful impact.
Responsibilities
- Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
- Developing data pipelines and integrating large-scale data, training, and inference systems
Requirements
- 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust
Skills
- Deep familiarity with modern ML frameworks such as JAX or PyTorch
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
- Comfortable with Linux systems and orchestration tools
- Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
- This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
- All employees are expected to be hands-on and to contribute directly to the company's mission.
- Work ethic and strong prioritization skills are important.
- All employees are expected to have strong communication skills.
- They should be able to concisely and accurately share knowledge with their teammates.
Compensation
- $180,000 - $440,000 USD
Benefits
- Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
- Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline
- 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
Equal opportunity
- equal opportunity employer.
This listing is sourced directly from xAI's careers page and normalized into a canonical job model.