Deepgram

Deepgram

Applied ML Engineer

USA | Remote · Senior

Sponsorship not specifiedDetected 13 days ago
PythonDistributed SystemsCloud PlatformsCI/CDMachine LearningDeep LearningPyTorchEmbedded SystemsResearch

About the role

  • Deepgram's voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency.
  • There is no organization in the world that understands voice better than Deepgram.
  • At Deepgram, we expect an AI-first mindset-AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance.

Responsibilities

  • Own the research-to-production pipeline: take research checkpoints and turn them into production models, defining the repeatable path from a working result to a deployed, monitored, scaled service.
  • Partner directly with research scientists to productionize new models - translating experimental training and evaluation code into robust, reproducible, well-tested workflows.
  • Build and extend the tooling and abstractions that let researchers and engineers move models through training, evaluation, packaging, and deployment with minimal friction and maximal reproducibility.
  • Design and own model release gates - automated evaluation, regression detection, and quality/latency/throughput checks that decide whether a model is ready to ship.
  • Optimize models and serving for production: efficient inference, batching, memory and latency tuning, and the profiling work that turns a research model into something that performs economically at scale.
  • Strengthen the build and delivery layer for models on our custom infrastructure, spanning our GPU compute and cloud environments, so that shipping a model is fast, safe, and observable.
  • Build the feedback loop: instrument production model behavior, surface what's working and what isn't, and feed it back to research to accelerate the next iteration.
  • Believe the last mile from research to production is the most important - and most underrated - problem in applied ML, and you want to own it.
  • Like working at the seam between research and engineering, fluent enough in ML to partner with scientists and rigorous enough in systems to ship at scale.

Requirements

  • Strong software engineering fundamentals, with proficiency in Python and experience writing production-quality, well-tested ML code.
  • Hands-on experience taking ML models from research or prototype stage into production at scale - not just training models, but shipping and operating them.
  • Familiarity with serving and inference optimization - latency, throughput, batching, and resource efficiency for production model workloads.
  • Comfort operating across distributed systems and GPU compute, whether in the cloud, on bare metal, or both.
  • Experience with the research-to-production handoff specifically - building the systems and conventions that let research and engineering iterate together quickly.
  • Experience designing automated model evaluation and release-gating systems, including regression detection across model versions.
  • Experience with inference optimization techniques (quantization, distillation, compilation, or runtime tuning) for production serving.

This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.