Nuancelabs

Nuancelabs

Member of Technical Staff — Model Optimization and Inference (Experienced)

Seattle, Washington · Staff+

H1B sponsorship available$250k-$350kDetected 40 days ago
PythonReactFull-Stack DevelopmentPyTorchLLMsLogisticsHRISResearch

About the role

  • The next problem is making it fast enough to actually use in a real-time conversation - and that gap is enormous.
  • A model that responds in 3 seconds is a demo.
  • A model that responds in under 500ms is a product.

Responsibilities

  • Own end-to-end inference optimization across our model stack - LLMs, audio models, and diffusion-based components
  • Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
  • Build internal tooling that makes optimization work faster and more rigorous - profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
  • Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
  • We believe diverse teams build better AI.

Requirements

  • Significant hands-on experience with LLM inference optimization - you've shipped work on KV caching, memory layout, attention kernels, or batching strategies in a production or high-traffic research context
  • Proven proficiency with inference serving frameworks - vLLM, SGLang, TensorRT-LLM, or similar - including going well beyond default configurations and adapting them to non-standard workloads
  • Experience optimizing diffusion model inference (latency reduction, step distillation, caching, or kernel-level work)
  • Familiarity with speculative decoding or other inference-time acceleration techniques
  • Hands-on experience with post-training quantization (GPTQ, AWQ, or similar) and a clear sense of quality/performance tradeoffs
  • Familiarity with multimodal or streaming inference architectures
  • Experience deploying real-time AI systems with hard latency SLAs

Nice to have

  • Strong Python and PyTorch skills
  • comfort reading and writing CUDA or Triton kernels is a significant plus

Skills

  • Do your best work with the best tools, including unlimited tokens.

Compensation

  • $250,000 - $350,000 base salary, plus meaningful equity.

Benefits

  • Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
  • Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
  • Commuter benefits: We help cover the cost of getting to the office.
  • We think long-term ownership matters and structure equity accordingly.

Company info

  • About Nuance Labs
  • Labs is building photorealistic, real-time AI avatars with emotional intelligence:
  • a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
  • We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
  • The team is small, the work is real, and the problems are unsolved.
  • How Nuance Differentiates
  • Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
  • Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
  • That's a 10x improvement, and it demands rethinking the entire stack.
  • What We're Looking For

Equal opportunity

  • equal opportunity employer.

Visa & Work Authorization

  • We sponsor visas (O-1, H-1B, green card) from day one.

This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.