Nuancelabs

Nuancelabs

Member of Technical Staff — ML Infra (Data)

Seattle, Washington · Staff+

H1B sponsorship available$200k-$300kDetected 39 days ago
ReactMachine LearningSparkData EngineeringLogisticsHRISResearch

About the role

  • Model quality is ultimately a data problem.
  • The best architecture and the best training run can't outrun bad, slow, or poorly curated data - and at the scale we're operating, the difference between a good data pipeline and a great one shows up directly in the model.
  • We're looking for someone who lives and breathes data at scale.

Responsibilities

  • Design, build, and operate large-scale data pipelines for ingestion, processing, filtering, and curation of multimodal training data (video, audio, text)
  • Optimize pipeline throughput and efficiency at scale
  • Build and maintain data quality systems - deduplication, filtering, validation, and quality scoring at scale
  • Manage petabyte-scale datasets: storage architecture, versioning, lineage tracking, and cost efficiency
  • Build tooling and infrastructure that makes the research team faster - efficient data access, reproducible processing, and fast iteration loops
  • Proven experience building and operating large-scale data pipelines in production - you've processed data at a scale where naive approaches break
  • Optimize pipeline throughput and efficiency at scale; identify and eliminate bottlenecks across compute, I/O, and storage
  • We believe diverse teams build better AI.
  • Experience building data pipelines for large-scale model training (pre-training or fine-tuning)

Requirements

  • Strong proficiency with distributed data processing frameworks - Spark, Ray, Dask, or similar - and a clear sense of when to use each
  • Ability to move fast: you can take a prototype script from a researcher and ship a production version in days, not weeks
  • Familiarity with data versioning and lineage tools (DVC, Delta Lake, Apache Iceberg, etc.)
  • Experience with streaming data pipelines or online data processing

Nice to have

  • Experience with multimodal data (video, audio) is a strong plus - understanding of formats, codecs, and processing libraries (FFmpeg, decord, etc.)
  • Familiarity with ML data pipelines specifically - understanding of how data quality and format affect model training

Skills

  • Do your best work with the best tools, including unlimited tokens.

Compensation

  • $200,000 - $300,000 base salary, plus meaningful equity.

Benefits

  • Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
  • Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
  • Commuter benefits: We help cover the cost of getting to the office.

Company info

  • About Nuance Labs
  • Labs is building photorealistic, real-time AI avatars with emotional intelligence:
  • a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
  • We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
  • The team is small, the work is real, and the problems are unsolved.
  • How Nuance Differentiates
  • Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
  • Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
  • That's a 10x improvement, and it demands rethinking the entire stack.
  • What We're Looking For

Equal opportunity

  • equal opportunity employer.

Visa & Work Authorization

  • We sponsor visas (O-1, H-1B, green card) from day one.

This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.