World Labs
Performance Engineer (Inference, Training & GPU)
San Francisco · Senior
Sponsorship not specified$200k-$300kDetected 5 days ago
PythonGoRustC++Distributed SystemsMachine LearningPyTorchExcelSystems EngineeringResearchCommunication
About the role
- We are looking for a Performance Engineer to make World Labs' models train and serve as fast as the hardware allows.
- Running large generative world models at scale is a novel systems problem.
- You will find the bottlenecks - in kernels, in the serving path, in the training loop, in how we use our GPUs - and eliminate them.
Responsibilities
- Optimize inference and serving end to end - latency, throughput, batching, caching, and scheduling - to serve our models efficiently at production scale.
- drive kernel fusion, memory- and bandwidth-bound optimization, and low-precision (FP8/INT8) execution.
- Optimize training throughput and GPU utilization: parallelism strategies, communication/compute overlap, mixed precision, and eliminating pipeline stalls.
- Build performance models, profiling workflows, and observability that make throughput, latency, cost, utilization, and their tradeoffs legible across the stack.
- Partner with researchers to productionize models for serving and to make experiments run faster and more reliably.
- We're building something bigger than any one person.
Requirements
- Working knowledge of ML framework internals (PyTorch and/or JAX
- Strong proficiency in Python, with the ability to drop into C++/CUDA (and Rust or Go) as the work demands.
- You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance.
- Distributed-systems breadth is welcome, but secondary.
Nice to have
- Experience at an AI lab or ML-native company, optimizing systems used directly by researchers and productionizing research code.
- Low-precision and numerics depth: FP8/INT8 quantization, mixed-precision, and detecting numerical regressions across hardware platforms.
- Distributed systems for large-scale training and inference - collective communication (NCCL), interconnects (NVLink), model and tensor parallelism, and fault tolerance.
- A strong plus, but not a substitute for the core skills above.
- Experience serving generative, diffusion, video, or 3D/spatial models - not just text LLMs.
- Multi-accelerator experience (GPU plus TPU or Trainium) and partnering with hardware vendors on accelerator capabilities.
- Fearless Innovator: We need people who thrive on challenges and aren't afraid to tackle the impossible.
- Resilient Builder: Impacting Large World Models isn't a sprint
Compensation
- $200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire
- Total Compensation
- Base salary plus equity awards
- Salary History
- We do not request or consider prior compensation in making offers
- Cal. Lab. Code §1197.5 (Equal Pay Act)
- $200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire; actual offer based on experience, skills, and qualifications)
Benefits
- Base salary plus equity awards
Company info
- Own numerical correctness across precision, kernel, and hardware changes - treating correctness as part of performance, not separate from it.
- You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary.
Equal opportunity
- World Labs is an equal opportunity employer.
- We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law.
- We welcome all qualified applicants and are committed to providing reasonable accommodations throughout the hiring process upon request.
- California Pay Transparency
Apply directly at World Labs →Create a free account for alerts like thisView World Labs immigration profile
This listing is sourced directly from World Labs's careers page and normalized into a canonical job model.