Phizenix

Phizenix

Site Reliability Engineering — AI Accelerator Infrastructure

Santa Clara, CA (3 Days Onsite)

Sponsorship not specified$155k-$235kDetected 15 days ago
PythonBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsibleCI/CDLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringMachine LearningLLMsIncident ResponseEmbedded SystemsElectrical EngineeringResearchCollaboration

About the role

  • We value humility and believe in direct communication.
  • Our team is inclusive, and our differing perspectives allow for better solutions.
  • This is a hands-on, high-ownership role.

Responsibilities

  • Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
  • Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up.
  • Support and operate high-speed interconnect environments - InfiniBand, RoCE, or high-speed Ethernet - in lab and colo settings.
  • Own IaC and configuration management (Terraform, Ansible) for your infrastructure domains - all provisioning and changes through code, not manual steps.
  • Develop networking automations for cluster interconnects, VLAN management, and lab network configurations.
  • Design and maintain monitoring dashboards, alerting, and SLIs (Prometheus/Grafana, DataDog) for your infrastructure domains-ensuring signal quality and actionable alerts.
  • Detect performance issues, recommend solutions, and implement fixes that permanently improve system reliability.
  • Document platform configurations, access procedures, and operational runbooks for customer environments.
  • Partner with the DevOps team to ensure infrastructure reliability supports CI/CD pipeline performance and developer experience.
  • Support and operate platform services used by external customers for hardware and software deployment collaboration

Requirements

  • Hands-on experience with colocation or on-premises server infrastructure - physical hardware, rack networking, and bare-metal provisioning.
  • IaC experience with Terraform and/or Ansible - writing and maintaining production configurations, not just running existing playbooks.
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience)
  • 5+ years in SRE, infrastructure engineering, or systems administration.
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration.

Nice to have

  • Experience operating customer-facing infrastructure or platform services with external reliability expectations.

Skills

  • Ready to come find your playground?
  • designs and manufactures purpose-built AI inference silicon.
  • responsible for the reliability, automation, and observability of the infrastructure that the company runs on.
  • colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.

Compensation

  • California Pay Range
  • $155,000 - $235,000 USD

Benefits

  • Go programming for SRE tooling - health-check services, exporters, or auto-remediation agents.
  • Build automation to eliminate toil: host lifecycle management, fleet health checks, auto-remediation workflows, and self-service tooling for engineering teams.

This listing is sourced directly from Phizenix's careers page and normalized into a canonical job model.