Radixark
Member of Technical Staff — Inference
Palo Alto, CA · Staff+
Sponsorship not specifiedDetected 5 hours ago
PythonRustC++Distributed SystemsGitAWSMachine LearningLLMsSystems EngineeringResearch
About the role
- RadixArk is seeking a Member of Technical Staff - Inference to push the limits of large-scale AI inference.
- You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs.
- This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.
Responsibilities
- Design and build large-scale inference systems for frontier AI models
- Optimize latency, throughput, and GPU utilization in production inference
- Develop and improve model serving architectures and runtimes
- Collaborate with kernel, compiler, and systems teams on performance optimization
- Drive reliability and scalability of inference infrastructure
- Build tooling for observability, profiling, and performance analysis
- Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
- We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs.
- Join us in building infrastructure that gives real leverage back to the AI community.
Requirements
- 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
- Strong expertise in large-scale inference systems for LLMs or generative models
- Experience optimizing latency- and throughput-critical production systems
- Strong knowledge of distributed systems and networking fundamentals
- Proficiency in Python, Rust, C++, or Go for production systems
- Experience profiling and optimizing compute-intensive workloads
- Experience with LLM serving stacks (SGLang, vLLM, TensorRT-LLM, etc.)
- Familiarity with CUDA, Triton, or custom kernel optimization
- Experience with batching, KV-cache management, and scheduling strategies
- Experience running inference at scale (1000+ GPUs)
Compensation
- We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements.
- Compensation depends on location, experience, and level.
Benefits
- We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements.
Company info
- We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training.
Equal opportunity
- RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Visa & Work Authorization
- t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more
Apply directly at Radixark →Create a free account for alerts like thisView Radixark immigration profile
This listing is sourced directly from Radixark's careers page and normalized into a canonical job model.