Deepgram
Embedded AI Engineer, On-Device Models
USA | Remote
Sponsorship not specifiedDetected 14 days ago
C++LinuxMachine LearningEmbedded SystemsSignal ProcessingResearchCommunication
About the role
- Deepgram's speech AI models are among the fastest and most accurate in the world - and the next wave of voice experiences won't live only in the cloud.
- They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware.
- Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.
Responsibilities
- Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
- Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
- Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
- Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
- Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.
Requirements
- Experience delivering production systems on resource-constrained hardware - embedded systems, mobile, edge AI, or small low-power devices.
- Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
- Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
- Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
- A strong understanding of hardware-software interaction - CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management - and how they affect inference performance.
- Experience with real-time audio processing on embedded platforms - DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.
- Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.
- Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.
- Experience with secure, robust on-device deployment practices - code signing, encrypted model storage, and safe update mechanisms.
Apply directly at Deepgram →Create a free account for alerts like thisView Deepgram immigration profile
This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.