Primeintellect

Primeintellect

Research Engineer - Distributed Training

San Francisco · Exec

Sponsorship not specified$150k-$350kDetected 14 days ago
Node.jsFull-Stack DevelopmentDatabricksDatadogMachine LearningPyTorchSystems EngineeringResearchCommunication

About the role

  • The next generation of AI companies, enterprises, and research teams do not just need more GPUs.
  • Its spans the full stack of training, deploying and continuously improving models - compute, large-scale RL, environments, sandboxes, evals, and deployment.

Responsibilities

  • the infrastructure frontier AI labs build internally, made available to every ambitious AI team.
  • If you're excited about building the systems foundation for frontier-scale training and open superintelligence, we'd love to hear from you.

Requirements

  • Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference.
  • Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.
  • Strong understanding of GPU architecture, profiling, and performance debugging.
  • Comfort working in a fast-moving environment with ambiguous problems and high ownership.
  • Experience writing or optimizing CUDA / Triton kernels.
  • Experience with compiler or runtime optimization for ML systems.
  • Experience working on RL training infrastructure, rollout systems, or asynchronous training pipelines.
  • Experience with multi-node GPU clusters and high-performance networking.

Compensation

  • Cash Compensation Range of $150-350k, plus equity incentives, aligning your success with the growth and impact of Prime Intellect.

Benefits

  • Cash Compensation Range of $150-350k, plus equity incentives, aligning your success with the growth and impact of Prime Intellect.
  • Flexible work arrangements, with the option to work remotely or in-person at our offices in San Francisco.
  • Quarterly team off-sites, hackathons, conferences and learning opportunities.

Company info

  • Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models.

Visa & Work Authorization

  • Visa sponsorship and relocation assistance for international candidates.

This listing is sourced directly from Primeintellect's careers page and normalized into a canonical job model.