Nuancelabs

Nuancelabs

Member of Technical Staff — RL Research (Experienced)

Seattle, Washington · Staff+

H1B sponsorship available$300k-$500kDetected 41 days ago
SwiftReactAlgorithmsMachine LearningLogisticsHRISFirewallResearchAdaptability

About the role

  • This posting is aimed at experienced researchers and engineers who've operated at a senior to senior-staff level at big tech or a leading research lab.
  • Everyone at Nuance is MTS - we don't run title ladders - but we're hiring people who have already done this work at scale.
  • This role is broader than a traditional RL algorithm role.

Responsibilities

  • Build Nuance's RL/post-training stack from 0→1: rollout generation, policy optimization, reward/reference model serving, data feedback loops, evaluation, checkpointing, observability, and debugging.
  • Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
  • Design the systems abstractions that connect research ideas to production-scale
  • Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, queueing, checkpoint overhead, and research iteration speed.
  • You will be expected to understand modern post-training methods and build the infrastructure needed to run them at scale.
  • You will build Nuance's RL/post-training stack from 0→1 and scale it from 1→10.
  • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
  • Build evaluation and feedback loops for omni behavior: turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
  • We believe diverse teams build better AI.

Requirements

  • Significant hands-on experience with RL, RLHF, RLAIF, post-training, alignment, or large-scale fine-tuning for modern foundation models.
  • A track record reasoning about model behavior and training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
  • Experience with large-scale training or inference systems, including rollout generation, model serving, batching, queueing, GPU utilization, checkpointing, and debugging.
  • Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
  • Experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
  • Experience scaling mixed training/inference workloads across large GPU clusters.
  • Experience with adjacent areas such as distributed pretraining, data infrastructure, inference serving, simulation, human/AI feedback collection, or evaluation infrastructure.

Skills

  • Do your best work with the best tools, including unlimited tokens.

Compensation

  • $300,000 - $500,000 base salary, plus meaningful equity. We think long-term ownership matters and structure equity accordingly.
  • Visa sponsorship: We sponsor visas (O-1, H-1B, green card) from day one.
  • AI-native tooling: Do your best work with the best tools, including unlimited tokens.

Benefits

  • Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
  • Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
  • Commuter benefits: We help cover the cost of getting to the office.

Company info

  • About Nuance Labs
  • Labs is building photorealistic, real-time AI avatars with emotional intelligence:
  • a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
  • We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
  • The team is small, the work is real, and the problems are unsolved.
  • How Nuance Differentiates
  • Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
  • Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
  • That's a 10x improvement, and it demands rethinking the entire stack.
  • We're looking for a deeply technical Member of Technical Staff to own RL and post-training for large-scale omni models.

Equal opportunity

  • equal opportunity employer.

Visa & Work Authorization

  • We sponsor visas (O-1, H-1B, green card) from day one.

This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.