Ifm Us

Ifm Us

Research Scientist - Distributed Machine Learning

Sunnyvale, CA

H1B sponsorship available$150k-$450kDetected 408 days ago
Node.jsAlgorithmsKubernetesMachine LearningDeep LearningPyTorchNumPyNLPResearchCommunicationCollaborationProblem Solving

About the role

  • Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters.
  • Turn mathematical ideas into fast production code
  • Prototype new optimizers or attention methods (like in PyTorch/NumPy/JAX orothers ).

Responsibilities

  • Experiment Infrastructure - Build reusable modules, logging, and metrics dashboards that speed up research cycles.
  • Collaboration - Document designs clearly, run post-mortems, and partner with global research teams.
  • Design ablation studies and statistical tests that validate-or refute-new ideas.
  • Mentor interns and junior engineers through clear async design docs and code reviews.
  • Research × Engineering blend - Move breakthrough papers into real systems and publish your own results.
  • We are a dedicated research lab for building, understanding, using, and risk-managing foundation models.
  • Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
  • Create robust launch scripts, resilient checkpoints, and job monitoring (e.g. NCCL/GLOO/GPU).
  • Lead mixed-precision training, push bf16, fp8, etc, into daily runs, track their accuracy-vs-speed gains, and be able to analyze numeric stability
  • Build logging, metrics, and other experiment-tracking tools for rapid iteration.

Requirements

  • Experience with distributed training at scale (100+ GPUs).
  • Proven multi-node GPU work (Slurm, K8s, or Ray) and NCCL/GLOO debugging.
  • Experience implementing optimization algorithms (e.g., SGD variants, Adam, second-order methods).
  • Ability to translate math and
  • Must-Haves
  • 5 + years combined industry or hands-on research experience with large-scale deep-learning training.
  • Led at least one large-scale transformer pre-training run
  • Strong software engineering skills on large ML codebases
  • Ownership of mixed- or low-precision paths ( bf16, fp8, 4-bit ) with accuracy validation.
  • Clear written communication (design docs, RFCs, post-mortems).
  • Nice-to-Haves
  • NeurIPS / ICML / ICLR papers or open-source contributions to major ML frameworks.
  • Background in numerical computing.

Nice to have

  • Expert PyTorch or JAX/Flax plus DeepSpeed, FSDP, Megatron-LM, or MosaicML Composer.

Skills

  • Build and scale distributed pre-training frameworks

Compensation

  • $150k-$450k

Benefits

  • Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability

Visa & Work Authorization

  • This position is eligible for visa sponsorship.

This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.