Nuancelabs
Member of Technical Staff — RL Research (Experienced)
Seattle, Washington · Staff+
H1B sponsorship available$300k-$500kDetected 41 days ago
SwiftReactAlgorithmsMachine LearningLogisticsHRISFirewallResearchAdaptability
About the role
- This posting is aimed at experienced researchers and engineers who've operated at a senior to senior-staff level at big tech or a leading research lab.
- Everyone at Nuance is MTS - we don't run title ladders - but we're hiring people who have already done this work at scale.
- This role is broader than a traditional RL algorithm role.
Responsibilities
- Build Nuance's RL/post-training stack from 0→1: rollout generation, policy optimization, reward/reference model serving, data feedback loops, evaluation, checkpointing, observability, and debugging.
- Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
- Design the systems abstractions that connect research ideas to production-scale
- Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, queueing, checkpoint overhead, and research iteration speed.
- You will be expected to understand modern post-training methods and build the infrastructure needed to run them at scale.
- You will build Nuance's RL/post-training stack from 0→1 and scale it from 1→10.
- Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
- Build evaluation and feedback loops for omni behavior: turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
- We believe diverse teams build better AI.
Requirements
- Significant hands-on experience with RL, RLHF, RLAIF, post-training, alignment, or large-scale fine-tuning for modern foundation models.
- A track record reasoning about model behavior and training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
- Experience with large-scale training or inference systems, including rollout generation, model serving, batching, queueing, GPU utilization, checkpointing, and debugging.
- Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
- Experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
- Experience scaling mixed training/inference workloads across large GPU clusters.
- Experience with adjacent areas such as distributed pretraining, data infrastructure, inference serving, simulation, human/AI feedback collection, or evaluation infrastructure.
Skills
- Do your best work with the best tools, including unlimited tokens.
Compensation
- $300,000 - $500,000 base salary, plus meaningful equity. We think long-term ownership matters and structure equity accordingly.
- Visa sponsorship: We sponsor visas (O-1, H-1B, green card) from day one.
- AI-native tooling: Do your best work with the best tools, including unlimited tokens.
Benefits
- Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
- Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
- Commuter benefits: We help cover the cost of getting to the office.
Company info
- About Nuance Labs
- Labs is building photorealistic, real-time AI avatars with emotional intelligence:
- a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
- We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
- The team is small, the work is real, and the problems are unsolved.
- How Nuance Differentiates
- Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
- Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
- That's a 10x improvement, and it demands rethinking the entire stack.
- We're looking for a deeply technical Member of Technical Staff to own RL and post-training for large-scale omni models.
Equal opportunity
- equal opportunity employer.
Visa & Work Authorization
- We sponsor visas (O-1, H-1B, green card) from day one.
Apply directly at Nuancelabs →Create a free account for alerts like thisView Nuancelabs immigration profile
This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.