Andromeda
HPC Architect
North America Remote / San Francisco, CA · Full-time
Sponsorship not specifiedDetected 7 hours ago
KubernetesLinuxSite Reliability EngineeringProcurementSystems EngineeringResearch
> stay_score
odds of building a lasting career here
25Risky
Cap-exempt (no lottery)0
Sponsors this role35
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →> community_outcomes
No reports yet — be the first to help the next applicant.
About the role
- You're the person who decides whether a provider's cluster is good enough to join the Andromeda network, and the person who helps them get there when it isn't.
- That means vetting prospective providers against our quality bar and then working alongside their engineers to bring their clusters onto the network cleanly.
- You'll work closely with our compute procurement team to identify and qualify new providers, and you'll be the standing technical relationship with the providers we already have.
Responsibilities
- Define the qualification bar itself. Build the acceptance test suite, benchmark methodology, and quality thresholds from scratch, formalizing what currently exists only as SRE tribal knowledge. Own and evolve these standards as the fleet and the market change.
- Partner with compute procurement to identify and qualify new providers. Perform technical due diligence during sourcing, and a clear read on how much remediation a candidate cluster needs before procurement commits.
- Build and maintain technical relationships with existing providers: you're the engineer their engineers call, and the one who spots architectural or quality drift before it becomes a customer problem.
- Deep HPC experience: you've designed, built, or operated GPU clusters at meaningful scale and understand what makes them fast, stable, and debuggable.
- The ability to write standards others can build against: precise, testable, and usable by a provider's engineering team without you in the room.
- Since then, we've been quietly building the systems, network, and orchestration layer that makes the world's AI infrastructure more accessible.
- Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most.
Requirements
- Experience with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent, and the Linux systems engineering underneath all of it.
- You can walk a provider's facility and know what questions to ask.
- Comfort with ambiguity.
- Experience inside a neocloud, hyperscaler, colocation, or data-center provider.
Nice to have
- Strong fabric knowledge, ideally both InfiniBand and RoCE.
- You can evaluate a topology, interpret fabric diagnostics, and identify why a fabric underperforms, not just that it does.
Compensation
- Competitive compensation: + meaningful equity
Benefits
- for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO
- We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
- Our long-term vision is to build the liquidity layer for global AI compute.
Company info
- We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
- What We're Looking For
- Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.
Equal opportunity
- equal opportunity employer.
Apply directly at Andromeda →Create a free account for alerts like thisView Andromeda immigration profile
This listing is sourced directly from Andromeda's careers page and normalized into a canonical job model.