Hark

Hark

Infrastructure, Speech

San Jose · Full-time

Sponsorship not specified$180k-$450kDetected 73 days ago
GoRustKubernetesCI/CDPlatform EngineeringKafkaWebSocketsMachine LearningData EngineeringAgentic AIIncident ResponseSystems EngineeringVoIP

About the role

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
  • While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.
  • Positioned at the nexus of systems engineering and speech AI, you will be accountable for the reliability, latency, and performance of the infrastructure supporting our live speech models.

Responsibilities

  • Lead the evolution of the end-to-end infrastructure powering Hark's speech-to-speech models, including streaming pipelines, session management, and fault tolerance.
  • Collaborate with speech ML researchers to identify latency bottlenecks and translate complex requirements into robust infrastructure enhancements.
  • Manage capacity planning, cost efficiency, and the hardware lifecycle for the global speech inference fleet.
  • Build internal tooling and platform abstractions to streamline the developer experience for teams operating on speech infrastructure.
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.

Requirements

  • Possess 5+ years of expertise in infrastructure, systems, or platform engineering, including a minimum of 2 years dedicated to real-time or low-latency environments.
  • Showcase advanced proficiency in at least one systems-level or infrastructure-centric programming language.
  • Proven experience with container orchestration, sophisticated job scheduling, and multi-tenant resource management.
  • Track record of technical ownership over production systems where high reliability and rigorous latency constraints are the standard.

Nice to have

  • Deep expertise in Kubernetes (K8s), with a focus on GPU-aware orchestration and the management of latency-sensitive workloads.
  • Proficiency in Pulumi or comparable modern Infrastructure as Code (IaC) frameworks.
  • Technical familiarity with speech model architectures-including ASR, TTS, and end-to-end speech-to-speech-and their unique inference characteristics.
  • Hands-on experience with streaming data pipelines and transport layers, such as Kafka, WebSockets, or custom audio protocols.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.

Compensation

  • The US base salary range for this full-time position is between $180,000 - $450,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components/benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • Oversee system health and incident response, defining critical SLOs for real-time speech workloads where performance and uptime are paramount.

Company info

  • agentic systems that interact naturally with people and the real world.
  • To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.
  • We are seeking a Member of Technical Staff, Infrastructure Speech to lead and scale the backbone of Hark's real-time speech-to-speech engine.
  • This is a high-impact technical role designed for someone who excels in low-latency distributed environments and approaches infrastructure with a product-driven mindset.

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.