Baseten

Baseten

Software Engineer - Voice AI (Inference Runtime)

San Francisco

Sponsorship not specifiedDetected 89 days ago
PythonFull-Stack DevelopmentCode ReviewDockerKubernetesMachine LearningPyTorchSparkA/B TestingProduct StrategyCustomer SupportResearchCommunicationCollaboration

About the role

  • You'll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more.
  • You'll make a meaningful impact on people's daily lives and help reshape these industries.
  • This is a high-impact, high-ownership role.

Responsibilities

  • Own and lead Voice AI product areas end-to-end - from architecture and system design through implementation, rollout, and long-term production operations.
  • Design, build, and operate real-time, large-scale, high-performance model serving systems for STT, TTS, and voice agent workloads for mission-critical customer deployments
  • Drive cross-team collaboration with sister engineering teams to solve full-stack technical problems, align on priorities, and coordinate end-to-end delivery across the product surface area
  • Mentor teammates through code reviews, design docs, and technical leadership.
  • We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital.
  • Join us and help build the platform engineers turn to to ship AI products.
  • Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs.
  • Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development.
  • Voice is becoming the internet's next interface, but a production-grade Voice AI system is "hard to build" https://greylock.com/greymatter/voice-agents-easy-to-use-hard-to-build/.

Requirements

  • Bachelor's degree or higher in Computer Science or related field
  • Proven track record owning production-grade real-time, large-scale systems where tail latency (p99) matters.

Nice to have

  • Experience implementing pipeline-level model runtime optimizations such as dynamic batching, async scheduling, or decode-side throughput improvements.
  • Experience with containerization and orchestration technologies (Docker, Kubernetes), service meshes, or distributed scheduling.
  • Familiarity with speech/audio ML models (STT, TTS, speech-to-speech)
  • Familiarity with model-serving runtimes (vLLM, TensorRT, ONNX).
  • Familiarity with systems-level performance profiling across host-device boundaries (e.g. PyTorch Profiler), diagnosing GPU utilization issues
  • Company-facilitated 401(k)
  • Python proficiency is a plus.
  • Proficient coding abilities in one or more popular programming or scripting languages

Compensation

  • Competitive compensation, including meaningful equity.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
  • By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.

Company info

  • pre-sales prototyping, technical discovery, or working directly with customers to ship solutions.
  • At Baseten, we are committed to fostering a diverse and inclusive workplace.

This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.