Hark
Infrastructure, Speech
San Jose · Full-time
Sponsorship not specified$180k-$450kDetected 73 days ago
GoRustKubernetesCI/CDPlatform EngineeringKafkaWebSocketsMachine LearningData EngineeringAgentic AIIncident ResponseSystems EngineeringVoIP
About the role
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.
- Positioned at the nexus of systems engineering and speech AI, you will be accountable for the reliability, latency, and performance of the infrastructure supporting our live speech models.
Responsibilities
- Lead the evolution of the end-to-end infrastructure powering Hark's speech-to-speech models, including streaming pipelines, session management, and fault tolerance.
- Collaborate with speech ML researchers to identify latency bottlenecks and translate complex requirements into robust infrastructure enhancements.
- Manage capacity planning, cost efficiency, and the hardware lifecycle for the global speech inference fleet.
- Build internal tooling and platform abstractions to streamline the developer experience for teams operating on speech infrastructure.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
Requirements
- Possess 5+ years of expertise in infrastructure, systems, or platform engineering, including a minimum of 2 years dedicated to real-time or low-latency environments.
- Showcase advanced proficiency in at least one systems-level or infrastructure-centric programming language.
- Proven experience with container orchestration, sophisticated job scheduling, and multi-tenant resource management.
- Track record of technical ownership over production systems where high reliability and rigorous latency constraints are the standard.
Nice to have
- Deep expertise in Kubernetes (K8s), with a focus on GPU-aware orchestration and the management of latency-sensitive workloads.
- Proficiency in Pulumi or comparable modern Infrastructure as Code (IaC) frameworks.
- Technical familiarity with speech model architectures-including ASR, TTS, and end-to-end speech-to-speech-and their unique inference characteristics.
- Hands-on experience with streaming data pipelines and transport layers, such as Kafka, WebSockets, or custom audio protocols.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
Compensation
- The US base salary range for this full-time position is between $180,000 - $450,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components/benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- Oversee system health and incident response, defining critical SLOs for real-time speech workloads where performance and uptime are paramount.
Company info
- agentic systems that interact naturally with people and the real world.
- To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.
- We are seeking a Member of Technical Staff, Infrastructure Speech to lead and scale the backbone of Hark's real-time speech-to-speech engine.
- This is a high-impact technical role designed for someone who excels in low-latency distributed environments and approaches infrastructure with a product-driven mindset.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.