Hark

Hark

Lead Audio ML Engineer

San Jose · Full-time

Sponsorship not specifiedDetected 23 days ago
Machine LearningDeep LearningTensorFlowPyTorchData EngineeringAgentic AICadenceSignal ProcessingResearch

About the role

  • You will work alongside our DSP, firmware, and product teams to turn audio model research into production features that ship at scale.

Responsibilities

  • Implement and train audio models for wake-word detection, voice activity detection, source separation, speech enhancement and similar audio
  • Build and maintain training data pipelines, evaluation harnesses, and re-training cadence across model families
  • Partner with DSP and firmware engineers to integrate models into the Hark Audio Engine and DSP runtime
  • Collaborate with hardware and acoustics teams to characterize the signal conditions models must operate under
  • Profile and optimize models on target platforms (DSP, NPU, CPU) and define accuracy and resource budgets per product
  • 3+ years of professional experience building and shipping audio or speech ML models
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.

Requirements

  • Comfort working across the full ML lifecycle: data, training, evaluation, deployment, and monitoring
  • Experience collaborating with DSP, firmware, and hardware engineers on resource-constrained systems

Nice to have

  • Background shipping voice-first or far-field audio products
  • Experience with on-device wake-word, ASR front-ends, or speech enhancement at production scale
  • Familiarity with model compression techniques such as quantization, pruning, and distillation
  • Familiarity with Qualcomm AI stacks or similar alternatives from other providers
  • Open-source contributions to audio ML projects
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Compensation

  • The US base salary range for this full-time position is between $120,000 and $300,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components and benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
  • Strong fluency in PyTorch or TensorFlow and modern audio deep learning toolchains
  • Bonus Qualifications

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.