Deepgram
Systems Architect AI/ML Infrastructure
USA | Remote · Senior
Sponsorship not specifiedDetected 107 days ago
AWSKubernetesMachine LearningData EngineeringProduct ManagementProduct StrategyProcurementEmbedded SystemsNetwork EngineeringResearchLeadershipCollaboration
About the role
- Deepgram's voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency.
- There is no organization in the world that understands voice better than Deepgram.
- At Deepgram, we expect an AI-first mindset-AI use and comfort aren't optional, they're core to how we operate, innovate, and measure performance.
Responsibilities
- Define and drive the end-to-end infrastructure architecture for Deepgram's AI/ML workloads across production inference and research training
- Architect compute orchestration systems that efficiently schedule and manage GPU and CPU workloads across heterogeneous infrastructure
- Lead capacity planning across all infrastructure dimensions, modeling growth and ensuring Deepgram can scale ahead of demand
- Drive cost optimization and FinOps practices, identifying opportunities to reduce infrastructure spend without compromising performance or reliability
- Design burstable, elastic training infrastructure that can scale up for large training runs and scale down to minimize idle cost
- Establish architectural standards, design review processes, and technical documentation practices for infrastructure decisions
- Collaborate with engineering leadership to align infrastructure strategy with product roadmap and business objectives
Requirements
- Design storage architectures that handle the massive datasets required for speech and audio ML -- from high-throughput training data pipelines to low-latency model serving
- 7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure
- Strong experience with compute orchestration using Kubernetes, and an understanding of how to schedule diverse workloads efficiently
- Hands-on experience with GPU infrastructure -- procurement considerations, cluster design, driver and runtime management
- Strong understanding of networking fundamentals as they relate to infrastructure architecture (see our Network Engineer role for the deep specialist)
- Experience with infrastructure modeling and simulation for capacity planning
Compensation
- Annual wellness stipend
Apply directly at Deepgram →Create a free account for alerts like thisView Deepgram immigration profile
This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.