Hark

Hark

On-Device Research Engineer

San Jose · Full-time

Sponsorship not specified$120k-$300kDetected 27 days ago
Machine LearningDeep LearningTensorFlowPyTorchNLPAgentic AISignal ProcessingResearch

About the role

  • We are looking for an On-Device Research Engineer to compress large audio and multimodal models into student models that meet the size, latency, and power budgets of our shipping hardware.
  • This role sits between training and production.
  • You will take teacher models from our research pipeline and produce student models that run on DSP, NPU, and microcontroller targets across our product line.

Responsibilities

  • Design and execute distillation strategies (response, feature, and self-distillation) to compress teacher models into deployable students
  • Build a reusable distillation and compression toolchain that the broader audio ML team can adopt across model families
  • Partner with the broader audio ML team on training pipelines and with the runtime team on deployment targets
  • Comfort working close to hardware and reasoning about compute, memory bandwidth, and power as design constraints
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.

Requirements

  • Track record of producing models that have shipped to constrained devices

Nice to have

  • Experience with Hexagon DSP, NPUs, Ambiq class MCUs, or similar
  • Experience with knowledge distillation at scale, including teacher-ensemble or multi-stage distillation
  • Familiarity with neural architecture search and hardware-aware NAS
  • Background shipping voice-first or far-field audio products
  • Contributions to open-source compression toolchains (TFLite, ONNX Runtime, AIMET, and similar)
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Compensation

  • The US base salary range for this full-time position is between $120,000 - $300,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components/benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
  • 3+ years of professional experience in model compression, distillation, quantization, or efficient deep learning
  • Bonus Qualifications

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.