xAI
Software Engineer - Training/Inference (C++)
Palo Alto, CA
Sponsorship not specified$180k-$440kDetected 2 days ago
RustC++Full-Stack DevelopmentCI/CDLLMsLoad BalancingResearchLeadershipCommunication
About the role
- This is a high-impact role where your work directly determines how fast and reliably users interact with Grok at massive scale
Responsibilities
- Architect and implement scalable distributed infrastructure for model serving (load balancing, auto-scaling, batch scheduling, global KV cache).
- Optimize latency and throughput of model inference under real production workloads.
- Build reliable, high-concurrency serving systems that serve billions of users with 100% uptime, 0% error rate, and excellent tail latency.
- Develop custom tools to trace, replay, and fix issues across the full stack - from orchestration down to GPU kernels.
- Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates.
- Accelerate research on scaling test-time compute, RL rollout, and model-hardware co-design for next-generation systems.
- As a Member of Technical Staff - Inference, you will design and optimize large-scale model serving systems end-to-end.
- You will own everything from distributed infrastructure (global KV cache, continuous batching, load balancing, auto-scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding, tail latency).
- Inference, you will design and optimize large-scale model serving systems end-to-end.
Requirements
- Experience with large-scale, high-concurrent production serving.
- Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).
- Experience with testing, benchmarking, and reliability of inference services.
- Experience designing and implementing CI/CD infrastructure for inference.
Skills
- This organization is for individuals who appreciate challenging themselves and thrive on curiosity.
- All employees are expected to be hands-on and to contribute directly to the company's mission.
- Work ethic and strong prioritization skills are important.
- All employees are expected to have strong communication skills.
- They should be able to concisely and accurately share knowledge with their teammates.
Compensation
- $180,000 - $440,000 USD
Company info
- We are building the high-performance inference platform that serves Grok to millions of users every day with lightning speed and perfect reliability.
Equal opportunity
- equal opportunity employer.
This listing is sourced directly from xAI's careers page and normalized into a canonical job model.