nscaleoperationsukltd
Senior Site Reliability Engineer -AI Infrastructure Operations
Houston; San Francisco; Seattle · Senior
Sponsorship not specified$170k-$265kDetected 1 day ago
PythonGoDistributed SystemsKubernetesLinuxSite Reliability EngineeringRESTEmbedded SystemsSystems Engineering
About the role
- You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time.
Responsibilities
- Own reliability for critical production services end to end
- mentor other SREs through design review, pairing, and
- Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted
- Lead the hardest incidents and the root causes nobody else can crack
- Build the tooling and automation that removes toil for the whole team, not just your own surface
Nice to have
- Familiarity with high-performance networking (InfiniBand, RDMA).
- On-Call and Pace
- A quick note on the shape of the job.
- This role sits close to production, so there is an on-call rotation,
- and some weeks are busier than others.
- As a senior on the team, you help set how that rotation runs
- and, more to the point, how we make it lighter over time.
- We share the load fairly, and we treat every
Skills
- About Nscale
- Nscale is the GPU cloud built for AI.
- We run high-performance, cost-efficient infrastructure for AI-native
- We move with urgency, we tell each
- other the truth, and everyone here stays close to the infrastructure that makes AI work.
- This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up
Compensation
- $170,000 - $265,000 USD. Actual compensation varies with skill set, experience, and location, and the
- The range below reflects the base salary for the position.
- Actual compensation may vary based on job-related factors such as skill set, experience, education, and location.
- In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs.
Benefits
- Competitive base plus equity, reviewed every 12 months.
- role may be eligible for bonus and equity.
- Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Company info
- Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change
- 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production
- Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,
- Deep command of Linux, networking, and distributed systems, plus the judgment to know where
- Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the
- Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to
Apply directly at nscaleoperationsukltd →Create a free account for alerts like thisView nscaleoperationsukltd immigration profile
This listing is sourced directly from nscaleoperationsukltd's careers page and normalized into a canonical job model.