Nuancelabs
Member of Technical Staff — RL Research (New PhD Grad)
Seattle, Washington · Staff+
H1B sponsorship available$250k-$350kDetected 41 days ago
SwiftReactAlgorithmsMachine LearningLogisticsHRISFirewallResearchAdaptability
About the role
- This posting is aimed at researchers who are completing - or have recently completed - a PhD and want to do their best work at a fast-moving frontier lab.
- This role is broader than a traditional RL algorithm role.
- The work spans RL method development, rollout generation, reward modeling, policy optimization, evaluation, data feedback loops, serving, observability, and distributed execution.
Responsibilities
- Build Nuance's RL/post-training stack from 0→1: rollout generation, policy optimization, reward/reference model serving, data feedback loops, evaluation, checkpointing, observability, and debugging.
- Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
- Design the systems abstractions that connect research ideas to production-scale
- Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, queueing, checkpoint overhead, and research iteration speed.
- You'll be expected to understand modern post-training methods and help build the infrastructure needed to run them at scale.
- You'll help build Nuance's RL/post-training stack from 0→1 and scale it from 1→10.
- Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
- Build evaluation and feedback loops for omni behavior: turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
- We believe diverse teams build better AI.
- Strong software engineering fundamentals and the appetite to build real systems, not just prototypes.
Requirements
- Ability to reason about model behavior and training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
- Hands-on experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
- Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
- Experience with adjacent areas such as distributed pretraining, data infrastructure, inference serving, simulation, human/AI feedback collection, or evaluation infrastructure.
Skills
- Do your best work with the best tools, including unlimited tokens.
Compensation
- $250,000 - $350,000 base salary, plus meaningful equity. We think long-term ownership matters and structure equity accordingly.
- Visa sponsorship: We sponsor visas (O-1, H-1B, green card) from day one.
- AI-native tooling: Do your best work with the best tools, including unlimited tokens.
Benefits
- Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
- Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
- Commuter benefits: We help cover the cost of getting to the office.
Company info
- About Nuance Labs
- Labs is building photorealistic, real-time AI avatars with emotional intelligence:
- a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
- We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
- The team is small, the work is real, and the problems are unsolved.
- How Nuance Differentiates
- Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
- Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
- That's a 10x improvement, and it demands rethinking the entire stack.
- We're looking for a deeply technical Member of Technical Staff to own RL and post-training for large-scale omni models.
Equal opportunity
- equal opportunity employer.
Visa & Work Authorization
- We sponsor visas (O-1, H-1B, green card) from day one.
Apply directly at Nuancelabs →Create a free account for alerts like thisView Nuancelabs immigration profile
This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.