Luma AI
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Redwood City, CA
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound (~15% per draw). A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →H-1B wage level
the lottery is wage-weighted — each level is one more entry
The bottom of this range ($30,000) is below the Level I prevailing wage of $100,422. An H-1B cannot be filed below the prevailing wage, so an offer at the floor of this band could not be sponsored as posted.
DOL prevailing wage, 2026-27 wage year · Computer and Information Research Scientists (15-1221) · San Francisco-Oakland-Fremont, CA. Wage level is derived by USCIS from the offered wage, occupation and worksite; the occupation shown is inferred from the job title.
Employer immigration record
from this employer's Department of Labor filings
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- RL is how Luma's models go from capable to useful.
- RL at scale is a full-loop systems problem: training, rollout generation, environment execution, and reward computation running concurrently across thousands of GPUs, all needing to stay fast, stable, and correct together.
- It fits someone who has lived this - post-trained LLMs with RL, built environments and verifiers, and debugged asynchronous rollout pipelines at scale.
Requirements
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.
- Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
- Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models.
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
- Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
Nice to have
- Running RL training across 100+ GPUs, including asynchronous or disaggregated trainer/rollout architectures.
- Containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads.
- Research contributions in RL for LLMs, or open-source contributions to RL training frameworks.
- Luma is an equal opportunity employer.
Compensation
- $30k-$60k
Benefits
- We believe multimodality is critical for intelligence - the next step beyond language models comes from vision.
Company info
- Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world.
This listing is sourced directly from Luma AI's careers page and normalized into a canonical job model.