Together AI
Senior Machine Learning Engineer, Voice AI
San Francisco · Senior · Full-time
Sponsorship not specified$200k-$260kDetected 75 days ago
PythonAlgorithmsPlatform EngineeringMachine LearningPyTorchLLMsElectrical EngineeringSignal ProcessingResearch
About the role
- Our Voice AI platform powers production-grade, real-time voice agents and applications - serving speech-to-text and text-to-speech models with best-in-class latency and reliability.
- This is a foundational hire on a small, high-impact team.
- Voice inference has unique challenges - streaming audio, tokenization, real-time latency budgets - that require dedicated ML engineering focus.
Responsibilities
- Optimize inference performance for voice models (STT, TTS, speech-to-speech) - targeting best-in-class TTFB, throughput, and GPU utilization across our curated model set.
- Build and maintain a voice model evaluation framework - measuring WER across accents, languages, and noise conditions for STT
- Collaborate with model partners to integrate and optimize their models (Cartesia, Deepgram, Rime, and others) running on Together's infrastructure.
- Build and maintain a voice model evaluation framework - measuring WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation accuracy for TTS.
- You'll profile GPU utilization, design batching strategies for streaming audio, and ensure new model architectures can go from research to production quickly.
- Own the model serving stack that powers Together's voice platform across STT, TTS, and speech-to-speech.
Requirements
- 5+ years of experience in ML engineering, with a focus on model serving, inference optimization, or ML infrastructure.
- Hands-on experience with LLM serving engines (vLLM, SGLang, TensorRT-LLM, or similar) - comfortable reading and modifying engine internals, not just using APIs.
- Strong proficiency in Python and PyTorch
- experience with GPU profiling and optimization (CUDA, memory management, kernel-level debugging).
- Track record of shipping ML systems to production with measurable performance improvements.
- Comfort working on a small, early-stage team where you'll wear multiple hats and move fast.
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
Nice to have
- Experience with speech and audio ML (ASR, TTS architectures, audio signal processing) is a strong plus but not required - you can learn this quickly if you have strong ML engineering fundamentals.
- Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC) is a plus.
- Experience training or fine-tuning speech models is a plus.
Compensation
- We offer competitive compensation, startup equity, health insurance and other competitive benefits.
- The US base salary range for this full-time position is: $200,000 - $260,000 + equity + benefits.
- Our salary ranges are determined by location, level and role.
- Individual compensation will be determined by experience, skills, and job-related knowledge.
Company info
- Contribute to voice model fine-tuning capabilities (STT and TTS) as we enable customers to build differentiated voice experiences on Together.
- We're looking for a Senior ML Engineer to drive the model serving layer for voice workloads.
Equal opportunity
- Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Visa & Work Authorization
- t opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more
Apply directly at Together AI →Create a free account for alerts like thisView Together AI immigration profile
This listing is sourced directly from Together AI's careers page and normalized into a canonical job model.