World Labs

World Labs

Performance Engineer (Inference, Training & GPU)

San Francisco · Senior

Sponsorship not specified$200k-$300kDetected 5 days ago
PythonGoRustC++Distributed SystemsMachine LearningPyTorchExcelSystems EngineeringResearchCommunication

About the role

  • We are looking for a Performance Engineer to make World Labs' models train and serve as fast as the hardware allows.
  • Running large generative world models at scale is a novel systems problem.
  • You will find the bottlenecks - in kernels, in the serving path, in the training loop, in how we use our GPUs - and eliminate them.

Responsibilities

  • Optimize inference and serving end to end - latency, throughput, batching, caching, and scheduling - to serve our models efficiently at production scale.
  • drive kernel fusion, memory- and bandwidth-bound optimization, and low-precision (FP8/INT8) execution.
  • Optimize training throughput and GPU utilization: parallelism strategies, communication/compute overlap, mixed precision, and eliminating pipeline stalls.
  • Build performance models, profiling workflows, and observability that make throughput, latency, cost, utilization, and their tradeoffs legible across the stack.
  • Partner with researchers to productionize models for serving and to make experiments run faster and more reliably.
  • We're building something bigger than any one person.

Requirements

  • Working knowledge of ML framework internals (PyTorch and/or JAX
  • Strong proficiency in Python, with the ability to drop into C++/CUDA (and Rust or Go) as the work demands.
  • You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance.
  • Distributed-systems breadth is welcome, but secondary.

Nice to have

  • Experience at an AI lab or ML-native company, optimizing systems used directly by researchers and productionizing research code.
  • Low-precision and numerics depth: FP8/INT8 quantization, mixed-precision, and detecting numerical regressions across hardware platforms.
  • Distributed systems for large-scale training and inference - collective communication (NCCL), interconnects (NVLink), model and tensor parallelism, and fault tolerance.
  • A strong plus, but not a substitute for the core skills above.
  • Experience serving generative, diffusion, video, or 3D/spatial models - not just text LLMs.
  • Multi-accelerator experience (GPU plus TPU or Trainium) and partnering with hardware vendors on accelerator capabilities.
  • Fearless Innovator: We need people who thrive on challenges and aren't afraid to tackle the impossible.
  • Resilient Builder: Impacting Large World Models isn't a sprint

Compensation

  • $200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire
  • Total Compensation
  • Base salary plus equity awards
  • Salary History
  • We do not request or consider prior compensation in making offers
  • Cal. Lab. Code §1197.5 (Equal Pay Act)
  • $200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire; actual offer based on experience, skills, and qualifications)

Benefits

  • Base salary plus equity awards

Company info

  • Own numerical correctness across precision, kernel, and hardware changes - treating correctness as part of performance, not separate from it.
  • You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary.

Equal opportunity

  • World Labs is an equal opportunity employer.
  • We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law.
  • We welcome all qualified applicants and are committed to providing reasonable accommodations throughout the hiring process upon request.
  • California Pay Transparency

This listing is sourced directly from World Labs's careers page and normalized into a canonical job model.