Deepgram

Deepgram

ML Ops Infrastructure Engineer

USA | Remote

Sponsorship not specifiedDetected 107 days ago
PythonDockerKubernetesTerraformCI/CDPrometheusGrafanaDatadogDevOpsMachine LearningMLOpsA/B TestingResearchCollaborationProblem Solving

About the role

  • Deepgram's voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency.
  • There is no organization in the world that understands voice better than Deepgram.
  • At Deepgram, we expect an AI-first mindset-AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance.

Responsibilities

  • Design and build CI/CD pipelines specifically tailored for ML model development, validation, and deployment
  • Architect and maintain model deployment pipelines that move models from research environments through staging to production with confidence
  • Build A/B testing infrastructure that enables controlled rollouts of new models and measures real-world performance impact
  • Implement comprehensive monitoring for model performance in production -- accuracy metrics, latency, drift detection, and regression alerts
  • Develop automated retraining pipelines that trigger on data changes, performance degradation, or scheduled cadences
  • Create and maintain build and test environments that mirror production, giving researchers high-fidelity feedback before deployment
  • Collaborate with research engineers to define and enforce model quality gates before production promotion
  • Optimize model serving infrastructure for latency, throughput, and cost efficiency

Requirements

  • 4+ years of experience in MLOps, DevOps, or infrastructure engineering with a focus on ML systems
  • Strong proficiency in Python and experience building automation and tooling for ML workflows
  • Deep experience with CI/CD systems and building pipelines for software and model delivery
  • Hands-on experience with Docker and Kubernetes for containerized workload management
  • Experience with model serving frameworks such as NVIDIA Triton Inference Server, TensorRT, or ONNX Runtime
  • Experience with Infrastructure as Code tools such as Terraform or Pulumi
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, or similar)
  • Experience with feature stores, data versioning, or ML metadata management

Compensation

  • Annual wellness stipend

Benefits

  • Build observability dashboards that give the team real-time insight into model health across all environments

Company info

  • Your work ensures that every model improvement our research team makes can be safely, quickly, and reliably delivered to the customers who depend on Deepgram's APIs for real-time voice AI.

This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.