Andromeda
Staff SRE, AI Infrastructure
North America Remote / San Francisco, CA · Staff+ · Full-time
Sponsorship not specifiedDetected 61 days ago
PythonGoNode.jsKubernetesTerraformAnsibleHelmLinuxSite Reliability EngineeringPyTorchLeadershipMentoring
About the role
- We're looking for someone with multiple years of hands-on experience operating GPU infrastructure at scale.
- You read NVIDIA release notes the day they drop.
- You have war stories about NCCL, fabric topology choices, and what it takes to keep a multi-thousand-GPU run healthy.
Responsibilities
- You lead the response, write the postmortem, and ship the systemic fix.
- You design with SLOs, error budgets, and failure modes in mind; they ship features; together you close the loop on every systemic issue.
- Partner with providers and DC teams on physical design - rack and pod layout, power and cooling envelopes, network topology, burn-in and validation - to keep failure modes out of production before they arrive.
- You build production tooling, controllers, and automation - not throwaway scripts.
- Hardware & Buildout Influence: Partner with providers and DC teams on physical design - rack and pod layout, power and cooling envelopes, network topology, burn-in and validation - to keep failure modes out of production before they arrive.
- Experience as the senior SRE partner in enterprise relationships for AI infrastructure or HPC.
Requirements
- Staff-Level SRE Track Record: A clear history of owning the reliability of load-bearing infrastructure.
- You understand memory hierarchies, ECC and SBE/DBE behavior, thermal envelopes, NVLink and NVSwitch topology, and hardware failure modes from direct production experience.
- High-Performance Networking, in Production: Real production experience with InfiniBand, RoCE, and NVLink fabrics for distributed training.
- Distributed Training Internals: Working knowledge of how large training jobs actually run - NCCL, CUDA, PyTorch distributed, FSDP, DeepSpeed, Megatron, and modern checkpointing/recovery patterns.
- When a 1,000+ GPU job stalls, you know where to look first.
- You can run an incident review with a customer's principal engineer, then walk into a deal review and frame the same content for a CTO buying compute.
Benefits
- Own the day-to-day health of thousands of GPUs across providers and generations.
- Build and own the telemetry, GPU health checks, fabric monitoring, and automated remediation that let us catch a degraded NVLink or a flaky transceiver before a customer does.
- Built or significantly contributed to a custom GPU health system, fleet manager, fabric controller, or on-call/incident tooling in production.
Company info
- Be the senior reliability voice in the room with sophisticated AI infra customers and providers.
- Run incident reviews with a customer's principal engineer.
- Scope demanding workloads.
- Sit in on architecture deep-dives and deal cycles where reliability credibility closes the room.
- Customer-Facing Technical Presence: Be the senior reliability voice in the room with sophisticated AI infra customers and providers.
Apply directly at Andromeda →Create a free account for alerts like thisView Andromeda immigration profile
This listing is sourced directly from Andromeda's careers page and normalized into a canonical job model.