Hark
Lead Audio ML Engineer
San Jose · Full-time
Sponsorship not specifiedDetected 23 days ago
Machine LearningDeep LearningTensorFlowPyTorchData EngineeringAgentic AICadenceSignal ProcessingResearch
About the role
- You will work alongside our DSP, firmware, and product teams to turn audio model research into production features that ship at scale.
Responsibilities
- Implement and train audio models for wake-word detection, voice activity detection, source separation, speech enhancement and similar audio
- Build and maintain training data pipelines, evaluation harnesses, and re-training cadence across model families
- Partner with DSP and firmware engineers to integrate models into the Hark Audio Engine and DSP runtime
- Collaborate with hardware and acoustics teams to characterize the signal conditions models must operate under
- Profile and optimize models on target platforms (DSP, NPU, CPU) and define accuracy and resource budgets per product
- 3+ years of professional experience building and shipping audio or speech ML models
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
Requirements
- Comfort working across the full ML lifecycle: data, training, evaluation, deployment, and monitoring
- Experience collaborating with DSP, firmware, and hardware engineers on resource-constrained systems
Nice to have
- Background shipping voice-first or far-field audio products
- Experience with on-device wake-word, ASR front-ends, or speech enhancement at production scale
- Familiarity with model compression techniques such as quantization, pruning, and distillation
- Familiarity with Qualcomm AI stacks or similar alternatives from other providers
- Open-source contributions to audio ML projects
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Compensation
- The US base salary range for this full-time position is between $120,000 and $300,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components and benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- Strong fluency in PyTorch or TensorFlow and modern audio deep learning toolchains
- Bonus Qualifications
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.