Hark
Infrastructure, Large-scale Training
San Jose · Full-time
Sponsorship not specified$180k-$450kDetected 83 days ago
Distributed SystemsKubernetesCI/CDPlatform EngineeringMachine LearningPyTorchAgentic AIIncident ResponseSystems EngineeringResearch
About the role
- You'll work at the intersection of systems engineering and machine learning infrastructure, owning the reliability, scalability, and efficiency of the compute platform that our research and engineering teams depend on.
- This is a high-impact, highly technical role suited for someone who thrives in complex distributed systems environments and cares deeply about infrastructure as a product.
Responsibilities
- Design, implement, and maintain Infrastructure as Code (IaC) best practices to enable repeatable, auditable, and scalable cluster provisioning.
- Own and evolve stable training infrastructure operating at the scale of 10,000+ GPUs, including job scheduling, fault tolerance, and network fabric optimization.
- Partner closely with ML researchers and engineers to understand compute bottlenecks and translate them into infrastructure improvements.
- Drive capacity planning, cost efficiency initiatives, and hardware lifecycle management across the GPU fleet.
- Hark is an artificial intelligence company building advanced, personalized intelligence.
- We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines.
Requirements
- 5+ years of experience in infrastructure, systems, or platform engineering, with at least 2 years working in ML or HPC environments.
- Strong proficiency in at least one systems or infrastructure programming language.
- Experience with container orchestration, job scheduling, and multi-tenant resource management.
- Proven track record owning production systems with high reliability requirements.
Nice to have
- Kubernetes (K8s) - particularly experience operating large, GPU-aware clusters.
- Pulumi or similar modern IaC tooling.
- Rust and/or Go for systems-level tooling and performance-critical services.
- Familiarity with PyTorch and Ray for understanding workload patterns and integration requirements.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- This information will be shared if an employment offer is extended.
- Deep understanding of networking fundamentals (RDMA, InfiniBand, or RoCE a plus) relevant to high-throughput training workloads.
Compensation
- The US base salary range for this full-time position is between $180,000 - $450,000 annually.
- The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience.
- The total compensation package may also include additional components and benefits depending on the specific role.
- This information will be shared if an employment offer is extended.
Benefits
- One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
- Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
Company info
- We are looking for a Member of Technical Staff, Infrastructure Compute to lead and manage large-scale GPU computing clusters powering our AI training and deployment workloads.
This listing is sourced directly from Hark's careers page and normalized into a canonical job model.