Forward
Incident Management Lead
Texas · Mid
Sponsorship not specifiedDetected 63 days ago
SQLSite Reliability EngineeringPlatform EngineeringIncident ResponseCadenceCommunication
About the role
- How fast you detect them, how quickly you act, and whether the organization actually learns from them - that is what separates payments companies that scale from ones that spiral.
- The question is not whether incidents will happen.
- It is whether Forward detects them in minutes or hours, resolves them with coordination or chaos, and fixes the root cause or patches the symptom.
Responsibilities
- Build the Detection Layer
- Build and maintain AI-assisted signal intelligence: use AIOps platforms (PagerDuty, Incident.io, or equivalent) to correlate alerts, suppress noise, and surface high-confidence incident precursors before they manifest as partner escalations.
- Own the governance loop over Support: review support ticket themes and escalation patterns on a weekly cadence to identify systemic issues before they cross into incident territory.
- Establish alert-to-noise discipline: define what a true signal looks like for each incident type, tune alerting thresholds, and drive Alert-to-Noise Ratio above 80% - the team acts on signals, not volume.
- Build and maintain the runbook library: pre-written, AI-augmented playbooks for the most common incident classes - submission failures, processing outages, ACH return spikes, TM system failures, compliance freezes - so the first 15 minutes of every incident are not spent figuring out who does what.
- Serve as the named incident owner when an incident is declared - responsible for coordinating Engineering, Support, GTM, and Operations from detection through resolution.
- Drive MTTD (Mean Time to Detect) and MTTA (Mean Time to Acknowledge) toward P1 targets: detection under 5 minutes, acknowledgment under 15 minutes for Sev-1 and Sev-2.
- Classify merchant and partner impact in real time: GPV-at-risk, number of affected merchants, partner SLA implications, and any regulatory reporting obligations under DORA or card network rules.
- A direct line to building the incident management function at a high-growth payments company from the ground up.
- Per-partner MTTR trends down quarter-over-quarter.
Requirements
- 4+ years in incident management, site reliability engineering, technical program management, or engineering operations - with direct ownership of production incident response.
- Hands-on experience with modern incident management platforms: PagerDuty, Incident.io, FireHydrant, Rootly, Blameless, or equivalent AIOps tooling.
- Familiarity with SLO/SLA frameworks, error budget concepts, and alert fatigue management.
- internal stakeholder updates, partner-facing status, and merchant-level communications where required - proactive, not reactive.
- Manage communications during incidents on a defined cadence: internal stakeholder updates, partner-facing status, and merchant-level communications where required - proactive, not reactive.
Nice to have
- Familiarity with DORA regulatory incident reporting obligations or equivalent financial services operational resilience frameworks.
- Experience integrating AI tooling into incident workflows: automated RCA, alert correlation, or runbook execution.
- SQL proficiency for pulling operational data to support incident diagnosis and post-incident analysis.
- Background in platform engineering or SRE at a payments or financial services company.
- What Success Looks Like
- Detection and Response
- MTTD (Mean Time to Detect): under 5 minutes for Sev-1, under 15 minutes for Sev-2.
- MTTA (Mean Time to Acknowledge): under 15 minutes for Sev-1, under 30 minutes for Sev-2.
Skills
- Declare incidents using a consistent severity framework (Sev-1 through Sev-3) with defined, documented SLAs for each tier.
- Drive
Compensation
- Competitive salary and equity package.
Benefits
- Competitive salary and equity package.
- Comprehensive health, dental, and vision benefits.
- Flexible work arrangements and generous PTO.
- Learning & development budget for conferences, courses, and certifications.
- SLO burn-rate alerting, synthetic health checks, deployment risk scoring, and real-time anomaly detection across submission, processing, and compliance pipelines.
- monthly scorecards showing incident frequency, MTTR, and resolution quality by partner - inputs into GTM conversations and partner health reviews.
Apply directly at Forward →Create a free account for alerts like thisView Forward immigration profile
This listing is sourced directly from Forward's careers page and normalized into a canonical job model.