Forward
Site Reliability Engineer
Santa Clara, CA
Sponsorship not specifiedDetected 6 days ago
PythonBashAWSGCPAzureCloud PlatformsKubernetesTerraformAnsiblePrometheusGrafanaDatadogDevOpsSite Reliability EngineeringIncident ResponseLoad TestingTCP/IPDNSFirewallLeadership
About the role
- This is not a "keep the lights on" SRE role.
- You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.
- If you thrive in environments where you're handed a problem rather than a playbook this role is for you.
Responsibilities
- Partner with engineering teams to embed reliability thinking into the SDLC - capacity planning, load testing, chaos engineering, and production readiness reviews
- Help define and build the SRE team as the company scales - this is a foundational hire with a path to leadership
- Drive the reliability and operational excellence of the Forward SaaS platform
Requirements
- 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
- Hands-on experience with Kubernetes and container orchestration in production environments
- Deep proficiency with observability tooling - Prometheus, Grafana, Datadog, Splunk, or similar
- Experience with cloud platforms - AWS, GCP, or Azure - including infrastructure as code (Terraform, Ansible, or equivalent)
- Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
- Ability to communicate clearly with both engineering teams and non-technical stakeholders - you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them
Nice to have
- Experience in a foundational or early SRE hire capacity at a growth stage company
- A siloed function - you will be deeply embedded with product and engineering teams
- A ticket-taker - you will be proactively identifying and solving reliability problems before they become incidents
- What This Role Is Not
- Strong fundamentals in networking - TCP/IP, DNS, routing, switching, firewalls, and load balancing.
- Experience with network management or observability platforms is a significant plus
Compensation
- According to IDC, Forward customers realize an average of $14.2 million in annual benefits through improved efficiency and security.
Company info
- Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the team always knows what's happening before customers do
Apply directly at Forward →Create a free account for alerts like thisView Forward immigration profile
This listing is sourced directly from Forward's careers page and normalized into a canonical job model.