Nuancelabs
Member of Technical Staff — Model Optimization and Inference (Experienced)
Seattle, Washington · Staff+
H1B sponsorship available$250k-$350kDetected 40 days ago
PythonReactFull-Stack DevelopmentPyTorchLLMsLogisticsHRISResearch
About the role
- The next problem is making it fast enough to actually use in a real-time conversation - and that gap is enormous.
- A model that responds in 3 seconds is a demo.
- A model that responds in under 500ms is a product.
Responsibilities
- Own end-to-end inference optimization across our model stack - LLMs, audio models, and diffusion-based components
- Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
- Build internal tooling that makes optimization work faster and more rigorous - profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
- Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
- We believe diverse teams build better AI.
Requirements
- Significant hands-on experience with LLM inference optimization - you've shipped work on KV caching, memory layout, attention kernels, or batching strategies in a production or high-traffic research context
- Proven proficiency with inference serving frameworks - vLLM, SGLang, TensorRT-LLM, or similar - including going well beyond default configurations and adapting them to non-standard workloads
- Experience optimizing diffusion model inference (latency reduction, step distillation, caching, or kernel-level work)
- Familiarity with speculative decoding or other inference-time acceleration techniques
- Hands-on experience with post-training quantization (GPTQ, AWQ, or similar) and a clear sense of quality/performance tradeoffs
- Familiarity with multimodal or streaming inference architectures
- Experience deploying real-time AI systems with hard latency SLAs
Nice to have
- Strong Python and PyTorch skills
- comfort reading and writing CUDA or Triton kernels is a significant plus
Skills
- Do your best work with the best tools, including unlimited tokens.
Compensation
- $250,000 - $350,000 base salary, plus meaningful equity.
Benefits
- Health: HSA plan with ~$2,000 in annual company contributions - roughly 2x what most big tech companies put in.
- Time off: 15 days of PTO plus public holidays, and we close the office for a full week at year-end.
- Commuter benefits: We help cover the cost of getting to the office.
- We think long-term ownership matters and structure equity accordingly.
Company info
- About Nuance Labs
- Labs is building photorealistic, real-time AI avatars with emotional intelligence:
- a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real person.
- We're a research company, with PhDs from MIT, UW, Oxford, CMU, and Johns Hopkins, and industry experience from Apple, Meta, Amazon AGI, and Discord.
- The team is small, the work is real, and the problems are unsolved.
- How Nuance Differentiates
- Most conversational AI avatars today are hacks - a face slapped on a speech-to-speech pipeline, stuck in the uncanny valley: emotionless, mechanical, one-turn-at-a-time.
- Current systems take 2-5 seconds to respond; natural conversation requires sub-500ms.
- That's a 10x improvement, and it demands rethinking the entire stack.
- What We're Looking For
Equal opportunity
- equal opportunity employer.
Visa & Work Authorization
- We sponsor visas (O-1, H-1B, green card) from day one.
Apply directly at Nuancelabs →Create a free account for alerts like thisView Nuancelabs immigration profile
This listing is sourced directly from Nuancelabs's careers page and normalized into a canonical job model.