Ifm Us
Research Scientist - Distributed Machine Learning
Sunnyvale, CA
H1B sponsorship available$150k-$450kDetected 408 days ago
Node.jsAlgorithmsKubernetesMachine LearningDeep LearningPyTorchNumPyNLPResearchCommunicationCollaborationProblem Solving
About the role
- Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters.
- Turn mathematical ideas into fast production code
- Prototype new optimizers or attention methods (like in PyTorch/NumPy/JAX orothers ).
Responsibilities
- Experiment Infrastructure - Build reusable modules, logging, and metrics dashboards that speed up research cycles.
- Collaboration - Document designs clearly, run post-mortems, and partner with global research teams.
- Design ablation studies and statistical tests that validate-or refute-new ideas.
- Mentor interns and junior engineers through clear async design docs and code reviews.
- Research × Engineering blend - Move breakthrough papers into real systems and publish your own results.
- We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
- Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
- Create robust launch scripts, resilient checkpoints, and job monitoring (e.g. NCCL/GLOO/GPU).
- Lead mixed-precision training, push bf16, fp8, etc, into daily runs, track their accuracy-vs-speed gains, and be able to analyze numeric stability
- Build logging, metrics, and other experiment-tracking tools for rapid iteration.
Requirements
- Experience with distributed training at scale (100+ GPUs).
- Proven multi-node GPU work (Slurm, K8s, or Ray) and NCCL/GLOO debugging.
- Experience implementing optimization algorithms (e.g., SGD variants, Adam, second-order methods).
- Ability to translate math and
- Must-Haves
- 5 + years combined industry or hands-on research experience with large-scale deep-learning training.
- Led at least one large-scale transformer pre-training run
- Strong software engineering skills on large ML codebases
- Ownership of mixed- or low-precision paths ( bf16, fp8, 4-bit ) with accuracy validation.
- Clear written communication (design docs, RFCs, post-mortems).
- Nice-to-Haves
- NeurIPS / ICML / ICLR papers or open-source contributions to major ML frameworks.
- Background in numerical computing.
Nice to have
- Expert PyTorch or JAX/Flax plus DeepSpeed, FSDP, Megatron-LM, or MosaicML Composer.
Skills
- Build and scale distributed pre-training frameworks
Compensation
- $150k-$450k
Benefits
- Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability
Visa & Work Authorization
- This position is eligible for visa sponsorship.
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.