Hark
On-Device Research Engineer
San Jose · Full-time
Sponsorship not specified$120k-$300kDetected 27 days ago
Machine LearningDeep LearningTensorFlowPyTorchNLPAgentic AISignal ProcessingResearch
About the role
- We are looking for an On-Device Research Engineer to compress large audio and multimodal models into student models that meet the size, latency, and power budgets of our shipping hardware.
- This role sits between training and production.
- You will take teacher models from our research pipeline and produce student models that run on DSP, NPU, and microcontroller targets across our product line.
Responsibilities
- Design and execute distillation strategies (response, feature, and self-distillation) to compress teacher models into deployable students
- Build a reusable distillation and compression toolchain that the broader audio ML team can adopt across model families
- Partner with the broader audio ML team on training pipelines and with the runtime team on deployment targets
- Comfort working close to hardware and reasoning about compute, memory bandwidth, and power as design constraints
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
Requirements
- Track record of producing models that have shipped to constrained devices
Nice to have
- Experience with Hexagon DSP, NPUs, Ambiq class MCUs, or similar
- Experience with knowledge distillation at scale, including teacher-ensemble or multi-stage distillation
- Familiarity with neural architecture search and hardware-aware NAS
- Background shipping voice-first or far-field audio products
- Contributions to open-source compression toolchains (TFLite, ONNX Runtime, AIMET, and similar)
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Compensation
- The US base salary range for this full-time position is between $120,000 - $300,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components/benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- 3+ years of professional experience in model compression, distillation, quantization, or efficient deep learning
- Bonus Qualifications
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.