Deepgram
Applied ML Engineer
USA | Remote · Senior
Sponsorship not specifiedDetected 13 days ago
PythonDistributed SystemsCloud PlatformsCI/CDMachine LearningDeep LearningPyTorchEmbedded SystemsResearch
About the role
- Deepgram's voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency.
- There is no organization in the world that understands voice better than Deepgram.
- At Deepgram, we expect an AI-first mindset-AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance.
Responsibilities
- Own the research-to-production pipeline: take research checkpoints and turn them into production models, defining the repeatable path from a working result to a deployed, monitored, scaled service.
- Partner directly with research scientists to productionize new models - translating experimental training and evaluation code into robust, reproducible, well-tested workflows.
- Build and extend the tooling and abstractions that let researchers and engineers move models through training, evaluation, packaging, and deployment with minimal friction and maximal reproducibility.
- Design and own model release gates - automated evaluation, regression detection, and quality/latency/throughput checks that decide whether a model is ready to ship.
- Optimize models and serving for production: efficient inference, batching, memory and latency tuning, and the profiling work that turns a research model into something that performs economically at scale.
- Strengthen the build and delivery layer for models on our custom infrastructure, spanning our GPU compute and cloud environments, so that shipping a model is fast, safe, and observable.
- Build the feedback loop: instrument production model behavior, surface what's working and what isn't, and feed it back to research to accelerate the next iteration.
- Believe the last mile from research to production is the most important - and most underrated - problem in applied ML, and you want to own it.
- Like working at the seam between research and engineering, fluent enough in ML to partner with scientists and rigorous enough in systems to ship at scale.
Requirements
- Strong software engineering fundamentals, with proficiency in Python and experience writing production-quality, well-tested ML code.
- Hands-on experience taking ML models from research or prototype stage into production at scale - not just training models, but shipping and operating them.
- Familiarity with serving and inference optimization - latency, throughput, batching, and resource efficiency for production model workloads.
- Comfort operating across distributed systems and GPU compute, whether in the cloud, on bare metal, or both.
- Experience with the research-to-production handoff specifically - building the systems and conventions that let research and engineering iterate together quickly.
- Experience designing automated model evaluation and release-gating systems, including regression detection across model versions.
- Experience with inference optimization techniques (quantization, distillation, compilation, or runtime tuning) for production serving.
Apply directly at Deepgram →Create a free account for alerts like thisView Deepgram immigration profile
This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.