Deepgram

Deepgram

Systems Architect AI/ML Infrastructure

USA | Remote · Senior

Sponsorship not specifiedDetected 107 days ago
AWSKubernetesMachine LearningData EngineeringProduct ManagementProduct StrategyProcurementEmbedded SystemsNetwork EngineeringResearchLeadershipCollaboration

About the role

  • Deepgram's voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency.
  • There is no organization in the world that understands voice better than Deepgram.
  • At Deepgram, we expect an AI-first mindset-AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance.

Responsibilities

  • Define and drive the end-to-end infrastructure architecture for Deepgram's AI/ML workloads across production inference and research training
  • Architect compute orchestration systems that efficiently schedule and manage GPU and CPU workloads across heterogeneous infrastructure
  • Lead capacity planning across all infrastructure dimensions, modeling growth and ensuring Deepgram can scale ahead of demand
  • Drive cost optimization and FinOps practices, identifying opportunities to reduce infrastructure spend without compromising performance or reliability
  • Design burstable, elastic training infrastructure that can scale up for large training runs and scale down to minimize idle cost
  • Establish architectural standards, design review processes, and technical documentation practices for infrastructure decisions
  • Collaborate with engineering leadership to align infrastructure strategy with product roadmap and business objectives

Requirements

  • Design storage architectures that handle the massive datasets required for speech and audio ML -- from high-throughput training data pipelines to low-latency model serving
  • 7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure
  • Strong experience with compute orchestration using Kubernetes, and an understanding of how to schedule diverse workloads efficiently
  • Hands-on experience with GPU infrastructure -- procurement considerations, cluster design, driver and runtime management
  • Strong understanding of networking fundamentals as they relate to infrastructure architecture (see our Network Engineer role for the deep specialist)
  • Experience with infrastructure modeling and simulation for capacity planning

Compensation

  • Annual wellness stipend

This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.