FAL

FAL

Software Engineer, Site Reliability

San Francisco

Sponsorship not specified$180k-$250kDetected 61 days ago
PythonBashKubernetesTerraformAnsibleCI/CDLinuxPrometheusGrafanaDatadogSite Reliability EngineeringMachine LearningSIEMEmbedded SystemsDNSBGP/OSPFLoad BalancingCommunication

About the role

  • You are a seasoned SRE who keeps production infrastructure running at scale.
  • You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.

Responsibilities

  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Build dashboards, alerting, and anomaly detection across our systems
  • Define and enforce SLOs and build out incident response processes
  • Manage and improve our networking, load balancing, and service mesh configurations
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)

Requirements

  • 5+ years experience in managing critical production systems and software development workflows
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
  • Proficiency in Python and either Go or Bash for tooling and automation
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)

Nice to have

  • Experience with managing GPU and AI/ML workloads
  • Experience with kernel-based monitoring and routing (eBPF, XDP)
  • Experience with security tooling (Falco, Coroot, SIEM)
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
  • Experience with distributed storage systems (Ceph, Longhorn, etc.)
  • Interesting and challenging work
  • Regular team events and offsites

Skills

  • fal is the generative media ecosystem powering the next generation of AI products.
  • About this role

Compensation

  • $180,000-250,000 plus equity + benefits (Range is based across 3 levels MId, Senior and Staff)
  • San Francisco, CA (willing to consider remote for Senior and Staff levels)
  • Interesting and challenging work
  • We are currently hiring in downtown San Francisco.
  • We offer relocation assistance to San Francisco.
  • Regular team events and offsites

Benefits

  • A lot of learning and growth opportunities
  • Health, dental, and vision insurance (US)

This listing is sourced directly from FAL's careers page and normalized into a canonical job model.