Hark

Hark

Audio Software Engineer

San Jose · Full-time

Sponsorship not specified$170k-$400kDetected 24 days ago
TypeScriptRustC++ReactMachine LearningAgentic AIDetection EngineeringSignal ProcessingVoIPResearch

About the role

  • We're hiring a Member of Technical Staff (Real-Time Audio) to join our Product Engineering team.
  • Hark's voice agent holds real-time, full-duplex conversations with people in homes, cars, and noisy rooms.
  • That experience is only as good as the audio underneath it.

Responsibilities

  • Own audio quality on the client: echo, self-interruption, dropouts, and clipping
  • Build and tune the browser audio pipeline with the Web Audio API, AudioWorklet, and getUserMedia constraints
  • Manage features end-to-end from prototyping through production
  • Collaborate with designers, platform engineers, and our speech team.
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
  • Owns features end-to-end and works comfortably in a shared production codebase.

Requirements

  • 5+ years of software engineering experience
  • Hands-on experience with WebRTC, AEC (echo cancellation), noise suppression, and VAD
  • C/C++ or Rust for production DSP, and experience shipping it to the browser via WebAssembly
  • Working knowledge of the browser audio stack: Web Audio API, AudioWorklet, and MediaStream constraints
  • Comfort with latency, buffering, and sample rates in a streaming audio pipeline

Nice to have

  • Experience working at a voice, speech, or video-conferencing company
  • ML for audio: noise suppression, VAD, or source separation (e.g. RNNoise, DeepFilterNet, Silero VAD), and on-device inference (ONNX Runtime, Core ML)
  • Familiarity with WebRTC internals (the Audio Processing Module, AEC3, Opus) and voice-agent frameworks (LiveKit, Pipecat)
  • TypeScript and React, and comfort working across the product frontend
  • Experience with target-speaker isolation, diarization, or barge-in and turn-detection systems for conversational AI.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Compensation

  • The US base salary range for this full-time position is between $170,000-$400,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components/benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

Company info

  • We're looking for someone who can do both: understand the signal processing and ship the code.

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.