Nuancelabs
Member of Technical Staff — ML Infra (Data)
Seattle, Washington · Staff+
H1B sponsorship available$200k-$300kDetected 39 days ago
ReactMachine LearningSparkData EngineeringLogisticsHRISResearch
About the role
- Model quality is ultimately a data problem.
- The best architecture and the best training run can't outrun bad, slow, or poorly curated data - and at the scale we're operating, the difference between a good data pipeline and a great one shows up directly in the model.
- We're looking for someone who lives and breathes data at scale.
Responsibilities
- Design, build, and operate large-scale data pipelines for ingestion, processing, filtering, and curation of multimodal training data (video, audio, text)
- Optimize pipeline throughput and efficiency at scale
- Build and maintain data quality systems - deduplication, filtering, validation, and quality scoring at scale
- Manage petabyte-scale datasets: storage architecture, versioning, lineage tracking, and cost efficiency
- Build tooling and infrastructure that makes the research team faster - efficient data access, reproducible processing, and fast iteration loops
- Proven experience building and operating large-scale data pipelines in production - you've processed data at a scale where naive approaches break
- Optimize pipeline throughput and efficiency at scale; identify and eliminate bottlenecks across compute, I/O, and storage
- We believe diverse teams build better AI.
- Experience building data pipelines for large-scale model training (pre-training or fine-tuning)
Requirements
- Strong proficiency with distributed data processing frameworks - Spark, Ray, Dask, or similar - and a clear sense of when to use each
- Ability to move fast: you can take a prototype script from a researcher and ship a production version in days, not weeks
- Familiarity with data versioning and lineage tools (DVC, Delta Lake, Apache Iceberg, etc.)
- Experience with streaming data pipelines or online data processing
Nice to have
- Experience with multimodal data (video, audio) is a strong plus - understanding of formats, codecs, and processing libraries (FFmpeg, decord, etc.)
- Familiarity with ML data pipelines specifically - understanding of how data quality and format affect model training
Skills
- Do your best work with the best tools, including unlimited tokens.
Compensation
- $200,000 - $300,000 base salary, plus meaningful equity.
Benefits
- Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
- Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
- Commuter benefits: We help cover the cost of getting to the office.
Company info
- About Nuance Labs
- Labs is building photorealistic, real-time AI avatars with emotional intelligence:
- a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
- We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
- The team is small, the work is real, and the problems are unsolved.
- How Nuance Differentiates
- Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
- Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
- That's a 10x improvement, and it demands rethinking the entire stack.
- What We're Looking For
Equal opportunity
- equal opportunity employer.
Visa & Work Authorization
- We sponsor visas (O-1, H-1B, green card) from day one.
Apply directly at Nuancelabs →Create a free account for alerts like thisView Nuancelabs immigration profile
This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.