Andromeda

Andromeda

Forward Deployed Engineer - SRE

North America Remote / San Francisco, CA · Full-time

Sponsorship not specifiedDetected 6 hours ago
PythonGoKubernetesTerraformAnsibleHelmLinuxSite Reliability EngineeringRESTMachine LearningIncident ResponseResearchCommunication

> stay_score

odds of building a lasting career here

25Risky
Cap-exempt (no lottery)0
Sponsors this role35
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70

Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.

Lottery odds assume a STEM candidate.

Personalize to your clock →

> community_outcomes

No reports yet — be the first to help the next applicant.

About the role

  • You will embed directly with the teams running large-scale training and inference on our clusters.
  • You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place.
  • Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform.

Responsibilities

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end
  • Own reliability outcomes for the accounts you're deployed on.
  • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it.
  • You will see our rough edges before anyone else does. Bring that signal back to influence the roadmap, file the hard bugs, and build the missing pieces yourself when that's the fastest path.
  • Working knowledge of how large training and inference jobs actually run. You don't need to design models, but you need to understand what's happening at the systems level when a large run stalls.
  • Since then, we've been quietly building the systems, network, and orchestration layer that makes the world's AI infrastructure more accessible.
  • Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most.

Requirements

  • Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent).
  • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training.
  • You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.
  • Working knowledge of how large training and inference jobs actually run.
  • Strong experience running Kubernetes in production with GPU workloads.
  • Experience with device plugins, topology-aware scheduling, multi-cluster, custom operators.
  • Experience with Slurm or other HPC schedulers is equally valued.
  • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).
  • You can go deep on architecture with a customer's infra team and clearly articulate tradeoffs to their leadership.
  • Proven track record leading incident response for complex distributed systems.

Compensation

  • Competitive compensation: + meaningful equity

Benefits

  • for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO
  • We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
  • Our long-term vision is to build the liquidity layer for global AI compute.

Company info

  • We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
  • We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework.
  • What We're Looking For - Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent).
  • What We're Looking For
  • Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

Equal opportunity

  • equal opportunity employer.

This listing is sourced directly from Andromeda's careers page and normalized into a canonical job model.