Forward

Forward

Site Reliability Engineer

Santa Clara, CA

Sponsorship not specifiedDetected 6 days ago
PythonBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsiblePrometheusGrafanaDatadogDevOpsSite Reliability EngineeringIncident ResponseLoad TestingTCP/IPDNSFirewallLeadership

About the role

  • This is not a "keep the lights on" SRE role.
  • You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.
  • If you thrive in environments where you're handed a problem rather than a playbook this role is for you.

Responsibilities

  • Partner with engineering teams to embed reliability thinking into the SDLC - capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales - this is a foundational hire with a path to leadership
  • Drive the reliability and operational excellence of the Forward SaaS platform

Requirements

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling - Prometheus, Grafana, Datadog, Splunk, or similar
  • Experience with cloud platforms - AWS, GCP, or Azure - including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders - you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

Nice to have

  • Experience in a foundational or early SRE hire capacity at a growth stage company
  • A siloed function - you will be deeply embedded with product and engineering teams
  • A ticket-taker - you will be proactively identifying and solving reliability problems before they become incidents
  • What This Role Is Not
  • Strong fundamentals in networking - TCP/IP, DNS, routing, switching, firewalls, and load balancing.
  • Experience with network management or observability platforms is a significant plus

Compensation

  • According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.

Company info

  • Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the team always knows what's happening before customers do

This listing is sourced directly from Forward's careers page and normalized into a canonical job model.