Deepgram

Deepgram

Embedded AI Engineer, On-Device Models

USA | Remote

Sponsorship not specifiedDetected 14 days ago
C++LinuxMachine LearningEmbedded SystemsSignal ProcessingResearchCommunication

About the role

  • Deepgram's speech AI models are among the fastest and most accurate in the world - and the next wave of voice experiences won't live only in the cloud.
  • They'll run directly on the small, low-power devices people carry, wear, and keep around their homes: phones, earbuds, wearables, appliances, cameras, and purpose-built consumer hardware.
  • Putting state-of-the-art speech models on devices with tight memory, compute, thermal, and battery budgets is a fundamentally different engineering problem, and it's one of the most important frontiers for bringing voice AI to everyone.

Responsibilities

  • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
  • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
  • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
  • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

Requirements

  • Experience delivering production systems on resource-constrained hardware - embedded systems, mobile, edge AI, or small low-power devices.
  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
  • A strong understanding of hardware-software interaction - CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management - and how they affect inference performance.
  • Experience with real-time audio processing on embedded platforms - DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.
  • Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.
  • Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.
  • Experience with secure, robust on-device deployment practices - code signing, encrypted model storage, and safe update mechanisms.

This listing is sourced directly from Deepgram's careers page and normalized into a canonical job model.