Hark

Hark

Infrastructure, Large-scale Training

San Jose · Full-time

Sponsorship not specified$180k-$450kDetected 83 days ago
Distributed SystemsKubernetesCI/CDPlatform EngineeringMachine LearningPyTorchAgentic AIIncident ResponseSystems EngineeringResearch

About the role

  • You'll work at the intersection of systems engineering and machine learning infrastructure, owning the reliability, scalability, and efficiency of the compute platform that our research and engineering teams depend on.
  • This is a high-impact, highly technical role suited for someone who thrives in complex distributed systems environments and cares deeply about infrastructure as a product.

Responsibilities

  • Design, implement, and maintain Infrastructure as Code (IaC) best practices to enable repeatable, auditable, and scalable cluster provisioning.
  • Own and evolve stable training infrastructure operating at the scale of 10,000+ GPUs, including job scheduling, fault tolerance, and network fabric optimization.
  • Partner closely with ML researchers and engineers to understand compute bottlenecks and translate them into infrastructure improvements.
  • Drive capacity planning, cost efficiency initiatives, and hardware lifecycle management across the GPU fleet.
  • Hark is an artificial intelligence company building advanced, personalized intelligence.
  • We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.

Requirements

  • 5+ years of experience in infrastructure, systems, or platform engineering, with at least 2 years working in ML or HPC environments.
  • Strong proficiency in at least one systems or infrastructure programming language.
  • Experience with container orchestration, job scheduling, and multi-tenant resource management.
  • Proven track record owning production systems with high reliability requirements.

Nice to have

  • Kubernetes (K8s) - particularly experience operating large, GPU-aware clusters.
  • Pulumi or similar modern IaC tooling.
  • Rust and/or Go for systems-level tooling and performance-critical services.
  • Familiarity with PyTorch and Ray for understanding workload patterns and integration requirements.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • This information will be shared if an employment offer is extended.
  • Deep understanding of networking fundamentals (RDMA, InfiniBand, or RoCE a plus) relevant to high-throughput training workloads.

Compensation

  • The US base salary range for this full-time position is between $180,000 - $450,000 annually.
  • The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
  • The total compensation package may also include additional components and benefits depending on the specific role.
  • This information will be shared if an employment offer is extended.

Benefits

  • One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
  • Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.

Company info

  • We are looking for a Member of Technical Staff, Infrastructure Compute to lead and manage large-scale GPU computing clusters powering our AI training and deployment workloads.

This listing is sourced directly from Hark's careers page and normalized into a canonical job model.