Ifm Us
Machine Learning Infrastructure Engineer
Sunnyvale, CA
Sponsorship not specified$150k-$450kDetected 369 days ago
PythonNode.jsAlgorithmsKubernetesMachine LearningDeep LearningPyTorchResearchProblem Solving
About the role
- We're looking for a distributed ML infrastructure engineer to help extend and scale our training systems.
Requirements
- While much of the work will support large-scale pre-training, pre-training experience is not required.
- Strong infrastructure and systems experience is what we value most.
- 5+ years of experience in ML systems, infra, or distributed training
- Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
- Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
- Familiarity with performance profiling, kernel fusion, or memory optimization
- Experience with large-scale pre-training
Compensation
- $150k-$450k
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.