Hark
Audio Software Engineer
San Jose · Full-time
Sponsorship not specified$170k-$400kDetected 24 days ago
TypeScriptRustC++ReactMachine LearningAgentic AIDetection EngineeringSignal ProcessingVoIPResearch
About the role
- We're hiring a Member of Technical Staff (Real-Time Audio) to join our Product Engineering team.
- Hark's voice agent holds real-time, full-duplex conversations with people in homes, cars, and noisy rooms.
- That experience is only as good as the audio underneath it.
Responsibilities
- Own audio quality on the client: echo, self-interruption, dropouts, and clipping
- Build and tune the browser audio pipeline with the Web Audio API, AudioWorklet, and getUserMedia constraints
- Manage features end-to-end from prototyping through production
- Collaborate with designers, platform engineers, and our speech team.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
- Owns features end-to-end and works comfortably in a shared production codebase.
Requirements
- 5+ years of software engineering experience
- Hands-on experience with WebRTC, AEC (echo cancellation), noise suppression, and VAD
- C/C++ or Rust for production DSP, and experience shipping it to the browser via WebAssembly
- Working knowledge of the browser audio stack: Web Audio API, AudioWorklet, and MediaStream constraints
- Comfort with latency, buffering, and sample rates in a streaming audio pipeline
Nice to have
- Experience working at a voice, speech, or video-conferencing company
- ML for audio: noise suppression, VAD, or source separation (e.g. RNNoise, DeepFilterNet, Silero VAD), and on-device inference (ONNX Runtime, Core ML)
- Familiarity with WebRTC internals (the Audio Processing Module, AEC3, Opus) and voice-agent frameworks (LiveKit, Pipecat)
- TypeScript and React, and comfort working across the product frontend
- Experience with target-speaker isolation, diarization, or barge-in and turn-detection systems for conversational AI.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Compensation
- The US base salary range for this full-time position is between $170,000-$400,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components/benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
Company info
- We're looking for someone who can do both: understand the signal processing and ship the code.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.