HUD

HUD

Platform Engineer

San Francisco · Full-time

Sponsorship not specifiedDetected 21 days ago
Backend DevelopmentAWSCloud PlatformsDockerKubernetesTerraformHelmCI/CDPlatform EngineeringMachine LearningAgentic AIIncident ResponseLogistics

About the role

  • We're looking for a Platform Engineer who can own the reliability, scale, performance, and developer experience of HUD's core infrastructure and backend systems.
  • This is not a pure infrastructure role. The right person has strong production infra experience, but also thinks like a backend engineer: they can reason about service architecture, queues, databases, APIs, deployment safety, performance bottlenecks, and how product requirements translate into resilient systems. You'll work across AWS, Kubernetes, Terraform, CI/CD, observability, and backend services to make HUD faster, more reliable, cheaper to run, and easier for engineers to build on.
  • This is not a pure infrastructure role.

Responsibilities

  • Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services
  • Build and maintain AWS infrastructure with Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management
  • Design and improve backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths
  • Build reliable CI/CD, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk
  • Write clean, maintainable code where needed to automate systems, improve backend services, and create internal tooling

Requirements

  • Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms
  • Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services
  • Experience reducing cloud spend through better architecture, autoscaling, workload placement, caching, cleanup systems, or observability

Nice to have

  • Have owned production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost
  • Have deep experience with AWS infrastructure and containerized systems
  • experience with tools like Terraform, Kubernetes/EKS, Docker, EC2, CodeBuild, ECR, S3, IAM, load balancers, networking, and secrets management is strongly preferred
  • Have built or operated CI/CD, environment management, release automation, observability, alerting, and incident response systems
  • Have strong backend engineering judgment and can reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes
  • Can write clean, maintainable code and apply strong software engineering judgment across product architecture, infrastructure, backend systems, and developer workflows

Compensation

  • Competitive compensation based on experience and location

Benefits

  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA
  • Company-wide holiday break (Christmas Eve to New Year's Day) on top of PTO and paid holidays
  • Other perks including an Equinox membership, 401k, and commuter benefits

Company info

  • We have 8 figures in funding and high revenue growth.
  • We're scaling profitably and quickly to meet very strong demand.
  • Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.

Visa & Work Authorization

  • Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US.

This listing is sourced directly from HUD's careers page and normalized into a canonical job model.