Radixark

Radixark

Member of Technical Staff — Training

Palo Alto, CA · Staff+

Sponsorship not specifiedDetected 7 hours ago
GitAWSMachine LearningLLMsResearch

About the role

  • You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, accuracy and reliability across 10k, or 100k+ of GPUs.
  • This role sits at the intersection of ML, systems, and performance engineering.
  • Your work will directly impact how next-generation AI models are trained and scaled.

Responsibilities

  • Optimize throughput, scalability, and hardware efficiency
  • Develop training frameworks and infrastructure tooling
  • Collaborate with model researchers to support frontier experiments
  • Build observability systems for training performance and reliability
  • Drive capacity planning and cluster utilization strategies
  • Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
  • We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs.

Requirements

  • 3+ years of experience in ML systems, or large-scale training infrastructure
  • Experience working on training / inference correctness or other precision-related problem
  • Experience debugging performance and stability issues in large post-training jobs
  • Experience improving training or inference efficiency.
  • Experience training 100+ billion-parameter models
  • Experience with train / inference optimization for large-scale RL or other production workload.
  • Familiarity with training stacks (e.g. Megatron-LM, FSDP, torchtitan, etc.) and inference stack (e.g. SGLang, vLLM, etc.)
  • Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL, AReaL, etc.)
  • Experience with RDMA, InfiniBand, NVLink, NCCL/RCCL, or high-speed GPU interconnects
  • Experience with checkpointing, fault recovery, and elastic training.

Compensation

  • We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements.
  • Compensation depends on location, experience, and level.

Company info

  • We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training.

Equal opportunity

  • RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Visa & Work Authorization

  • t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more

This listing is sourced directly from Radixark's careers page and normalized into a canonical job model.