FAL
Software Engineer, Site Reliability
San Francisco
Sponsorship not specified$180k-$250kDetected 61 days ago
PythonBashKubernetesTerraformAnsibleCI/CDLinuxPrometheusGrafanaDatadogSite Reliability EngineeringMachine LearningSIEMEmbedded SystemsDNSBGP/OSPFLoad BalancingCommunication
About the role
- You are a seasoned SRE who keeps production infrastructure running at scale.
- You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.
Responsibilities
- Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
- Build and maintain CI/CD pipelines and deployment infrastructure
- Build dashboards, alerting, and anomaly detection across our systems
- Define and enforce SLOs and build out incident response processes
- Manage and improve our networking, load balancing, and service mesh configurations
- Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
- Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
Requirements
- 5+ years experience in managing critical production systems and software development workflows
- Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
- Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
- Proficiency in Python and either Go or Bash for tooling and automation
- Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
Nice to have
- Experience with managing GPU and AI/ML workloads
- Experience with kernel-based monitoring and routing (eBPF, XDP)
- Experience with security tooling (Falco, Coroot, SIEM)
- Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
- Experience with distributed storage systems (Ceph, Longhorn, etc.)
- Interesting and challenging work
- Regular team events and offsites
Skills
- fal is the generative media ecosystem powering the next generation of AI products.
- About this role
Compensation
- $180,000-250,000 plus equity + benefits (Range is based across 3 levels MId, Senior and Staff)
- San Francisco, CA (willing to consider remote for Senior and Staff levels)
- Interesting and challenging work
- We are currently hiring in downtown San Francisco.
- We offer relocation assistance to San Francisco.
- Regular team events and offsites
Benefits
- A lot of learning and growth opportunities
- Health, dental, and vision insurance (US)
This listing is sourced directly from FAL's careers page and normalized into a canonical job model.