Radixark

Radixark

Member of Technical Staff — Cluster / Platform

Palo Alto, CA · Staff+

Sponsorship not specifiedDetected 8 hours ago
PythonGoRustC++Distributed SystemsGitAWSKubernetesLinuxMachine LearningLLMsSystems EngineeringResearch

About the role

  • RadixArk is looking for a Member of Technical Staff (Cluster / Platform) to architect and scale the core compute platform that powers frontier-level AI training and inference.
  • This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization.
  • Your work will directly impact how efficiently frontier AI models are trained and served.

Responsibilities

  • Design cluster management, scheduling, and resource allocation systems
  • Optimize performance, utilization, and reliability of GPU/TPU clusters
  • Drive observability, monitoring, and performance profiling for cluster infrastructure
  • Collaborate with ML and systems engineers to support frontier AI workloads
  • Lead capacity planning and infrastructure scaling strategies
  • Build internal platforms and tooling to improve developer productivity
  • Document architecture, operational practices, and reliability strategies
  • Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Requirements

  • 5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms
  • Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers)
  • Hands-on experience with GPU/TPU infrastructure in production environments
  • Proficiency in Go, Rust, C++, or Python for production systems
  • Experience debugging complex multi-layer issues across hardware, OS, networking, and distributed services

Nice to have

  • Experience with large-scale ML/AI workloads
  • Familiarity with RDMA, InfiniBand, or high-performance networking
  • Experience operating clusters at 1000+ GPU scale
  • Background in HPC or performance-critical systems
  • Open-source contributions in systems or infrastructure

Compensation

  • We offer competitive compensation with equity, comprehensive health benefits, and flexible work arrangements.
  • Compensation is determined by location, level, and experience.

Benefits

  • Contribute to long-term platform vision and technical direction

Company info

  • We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training.

Equal opportunity

  • RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Visa & Work Authorization

  • t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more

This listing is sourced directly from Radixark's careers page and normalized into a canonical job model.