Luma AI

Luma AI

Tech Lead Manager, Inference

Redwood City, CA · Senior

Sponsorship not specified$30k-$60kDetected 43 days ago
PythonC++Node.jsDistributed SystemsKubernetesMachine LearningPyTorchLLMsIncident ResponseResearchLeadershipMentoring

Stay score

odds of building a lasting career here

35Risky
Cap-exempt (no lottery)0
Sponsors this role74
Entry-level history0
PERM / green-card track0
Lottery odds (Level I)39
Fits your clock70

Thin sponsorship signal and lottery-bound (~15% per draw). A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.

Lottery odds assume a STEM candidate.

Personalize to your clock →

Employer immigration record

from this employer's Department of Labor filings

Files H-1B transfers

16 transfer filings in the last year, covering 16 workers. Median labor-condition decision: 7 days. An employer that already files transfers is one that can take over an existing H-1B.

Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.

Community outcomes

No reports yet — be the first to help the next applicant.

About the role

  • It's leadership by shipping: at least half your time stays hands-on in the serving stack, alongside hiring, growing the team, and setting technical direction.
  • If you want a hands-off management seat, this isn't it.
  • We believe multimodality is critical for intelligence - the next step beyond language models comes from vision.

Requirements

  • Strong Python and PyTorch, Kubernetes at scale, and experience with queues, scheduling, traffic control, and fleet management.
  • 8+ years in large-scale distributed systems or ML infrastructure, with several years building and operating model-serving or inference platforms in production.

Nice to have

  • Experience serving diffusion, video, or other multimodal generative models, and with FFmpeg/multimedia processing.
  • Modern networking stacks - RDMA (RoCE, InfiniBand), NVLink - and multi-node serving topologies.
  • Experience across heterogeneous accelerators (NVIDIA, AMD, TPU, Trainium) and the porting and validation that comes with them.
  • Contributions to open-source serving infrastructure (vLLM, SGLang, Ray, Kubernetes ecosystem).
  • Systems-language depth (Rust, C++, CUDA/HIP) for kernel- and runtime-level optimization.
  • Luma is an equal opportunity employer.
  • Experience running inference platforms at the thousands-of-GPUs scale across multiple clusters or clouds, and knowing what breaks there.
  • Technical leadership experience through rapid growth, with a genuine desire to stay at least half hands-on.

Compensation

  • $30k-$60k

Company info

  • Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world.

This listing is sourced directly from Luma AI's careers page and normalized into a canonical job model.