Phizenix
Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, CA (3 Days Onsite)
Sponsorship not specified$155k-$235kDetected 15 days ago
PythonBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringMachine LearningLLMsIncident ResponseEmbedded SystemsElectrical EngineeringResearchCollaboration
About the role
- We value humility and believe in direct communication.
- Our team is inclusive, and our differing perspectives allow for better solutions.
- This is a hands-on, high-ownership role.
Responsibilities
- Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
- Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up.
- Support and operate high-speed interconnect environments - InfiniBand, RoCE, or high-speed Ethernet - in lab and colo settings.
- Own IaC and configuration management (Terraform, Ansible) for your infrastructure domains - all provisioning and changes through code, not manual steps.
- Develop networking automations for cluster interconnects, VLAN management, and lab network configurations.
- Design and maintain monitoring dashboards, alerting, and SLIs (Prometheus/Grafana, DataDog) for your infrastructure domains-ensuring signal quality and actionable alerts.
- Detect performance issues, recommend solutions, and implement fixes that permanently improve system reliability.
- Document platform configurations, access procedures, and operational runbooks for customer environments.
- Partner with the DevOps team to ensure infrastructure reliability supports CI/CD pipeline performance and developer experience.
- Support and operate platform services used by external customers for hardware and software deployment collaboration
Requirements
- Hands-on experience with colocation or on-premises server infrastructure - physical hardware, rack networking, and bare-metal provisioning.
- IaC experience with Terraform and/or Ansible - writing and maintaining production configurations, not just running existing playbooks.
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience)
- 5+ years in SRE, infrastructure engineering, or systems administration.
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration.
Nice to have
- Experience operating customer-facing infrastructure or platform services with external reliability expectations.
Skills
- Ready to come find your playground?
- designs and manufactures purpose-built AI inference silicon.
- responsible for the reliability, automation, and observability of the infrastructure that the company runs on.
- colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
Compensation
- California Pay Range
- $155,000 - $235,000 USD
Benefits
- Go programming for SRE tooling - health-check services, exporters, or auto-remediation agents.
- Build automation to eliminate toil: host lifecycle management, fleet health checks, auto-remediation workflows, and self-service tooling for engineering teams.
Apply directly at Phizenix →Create a free account for alerts like thisView Phizenix immigration profile
This listing is sourced directly from Phizenix's careers page and normalized into a canonical job model.