Phizenix
Director, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, CA (3 Days Onsite) · Director
Sponsorship not specified$195k-$285kDetected 15 days ago
PythonAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringElectrical EngineeringTCP/IPLeadershipCommunicationCollaborationWriting
About the role
- You will build and lead Our Client's Site Reliability Engineering function from the ground up - owning the infrastructure that development, validation, and customer-facing deployments run on.
- This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate with our Client on hardware and software deployments.
- You are a hands-on engineering leader.
Responsibilities
- Leadership & Organizational Build-Out
- Own the SRE function end-to-end: define the team's charter, establish SRE as a discipline within our client's engineering culture, and drive buy-in across hardware, software, and executive stakeholders who have operated without a dedicated SRE team.
- Hire, develop, and retain a team of 3-5 SRE engineers
- Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery
- Reliability & Observability - Building From Scratch
- Establish SRE process from a zero baseline: define SLIs and SLOs, build error budgets, design on-call rotations, and create the incident management framework our client currently lacks.
- Own 24×7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services - designing for failure domains, progressive delivery, and strict change control at every tier.
- Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery; the two functions must operate as a unified platform.
- Ability to operate in a high-ambiguity, low-process environment - you build the structure, you don't inherit it.
Skills
- AWS, Azure, GCP alongside on-prem/colo - unified observability and IaC across all tiers.
Compensation
- California Pay Range
Benefits
- Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
Company info
- Design and scale a follow-the-sun on-call model as our client expands globally; the framework you build now will be the foundation the team inherits.
- Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our client's accelerator roadmap.
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering.
- 5+ years leading SRE or infrastructure engineering teams - including experience building or significantly rebuilding a function, not just managing a steady-state team.
- Demonstrated track record of establishing
Apply directly at Phizenix →Create a free account for alerts like thisView Phizenix immigration profile
This listing is sourced directly from Phizenix's careers page and normalized into a canonical job model.