Phizenix

Phizenix

Director, Site Reliability Engineering — AI Accelerator Infrastructure

Santa Clara, CA (3 Days Onsite) · Director

Sponsorship not specified$195k-$285kDetected 15 days ago
PythonAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringElectrical EngineeringTCP/IPLeadershipCommunicationCollaborationWriting

About the role

  • You will build and lead Our Client's Site Reliability Engineering function from the ground up - owning the infrastructure that development, validation, and customer-facing deployments run on.
  • This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate with our Client on hardware and software deployments.
  • You are a hands-on engineering leader.

Responsibilities

  • Leadership & Organizational Build-Out
  • Own the SRE function end-to-end: define the team's charter, establish SRE as a discipline within our client's engineering culture, and drive buy-in across hardware, software, and executive stakeholders who have operated without a dedicated SRE team.
  • Hire, develop, and retain a team of 3-5 SRE engineers
  • Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery
  • Reliability & Observability - Building From Scratch
  • Establish SRE process from a zero baseline: define SLIs and SLOs, build error budgets, design on-call rotations, and create the incident management framework our client currently lacks.
  • Own 24×7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services - designing for failure domains, progressive delivery, and strict change control at every tier.
  • Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery; the two functions must operate as a unified platform.
  • Ability to operate in a high-ambiguity, low-process environment - you build the structure, you don't inherit it.

Skills

  • AWS, Azure, GCP alongside on-prem/colo - unified observability and IaC across all tiers.

Compensation

  • California Pay Range

Benefits

  • Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.

Company info

  • Design and scale a follow-the-sun on-call model as our client expands globally; the framework you build now will be the foundation the team inherits.
  • Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our client's accelerator roadmap.
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering.
  • 5+ years leading SRE or infrastructure engineering teams - including experience building or significantly rebuilding a function, not just managing a steady-state team.
  • Demonstrated track record of establishing

This listing is sourced directly from Phizenix's careers page and normalized into a canonical job model.