Andromeda

Andromeda

Staff SRE, AI Infrastructure

North America Remote / San Francisco, CA · Staff+ · Full-time

Sponsorship not specifiedDetected 61 days ago
PythonGoNode.jsKubernetesTerraformAnsibleHelmLinuxSite Reliability EngineeringPyTorchLeadershipMentoring

About the role

  • We're looking for someone with multiple years of hands-on experience operating GPU infrastructure at scale.
  • You read NVIDIA release notes the day they drop.
  • You have war stories about NCCL, fabric topology choices, and what it takes to keep a multi-thousand-GPU run healthy.

Responsibilities

  • You lead the response, write the postmortem, and ship the systemic fix.
  • You design with SLOs, error budgets, and failure modes in mind; they ship features; together you close the loop on every systemic issue.
  • Partner with providers and DC teams on physical design - rack and pod layout, power and cooling envelopes, network topology, burn-in and validation - to keep failure modes out of production before they arrive.
  • You build production tooling, controllers, and automation - not throwaway scripts.
  • Hardware & Buildout Influence: Partner with providers and DC teams on physical design - rack and pod layout, power and cooling envelopes, network topology, burn-in and validation - to keep failure modes out of production before they arrive.
  • Experience as the senior SRE partner in enterprise relationships for AI infrastructure or HPC.

Requirements

  • Staff-Level SRE Track Record: A clear history of owning the reliability of load-bearing infrastructure.
  • You understand memory hierarchies, ECC and SBE/DBE behavior, thermal envelopes, NVLink and NVSwitch topology, and hardware failure modes from direct production experience.
  • High-Performance Networking, in Production: Real production experience with InfiniBand, RoCE, and NVLink fabrics for distributed training.
  • Distributed Training Internals: Working knowledge of how large training jobs actually run - NCCL, CUDA, PyTorch distributed, FSDP, DeepSpeed, Megatron, and modern checkpointing/recovery patterns.
  • When a 1,000+ GPU job stalls, you know where to look first.
  • You can run an incident review with a customer's principal engineer, then walk into a deal review and frame the same content for a CTO buying compute.

Benefits

  • Own the day-to-day health of thousands of GPUs across providers and generations.
  • Build and own the telemetry, GPU health checks, fabric monitoring, and automated remediation that let us catch a degraded NVLink or a flaky transceiver before a customer does.
  • Built or significantly contributed to a custom GPU health system, fleet manager, fabric controller, or on-call/incident tooling in production.

Company info

  • Be the senior reliability voice in the room with sophisticated AI infra customers and providers.
  • Run incident reviews with a customer's principal engineer.
  • Scope demanding workloads.
  • Sit in on architecture deep-dives and deal cycles where reliability credibility closes the room.
  • Customer-Facing Technical Presence: Be the senior reliability voice in the room with sophisticated AI infra customers and providers.

This listing is sourced directly from Andromeda's careers page and normalized into a canonical job model.