Forward

Forward

Incident Management Lead

Texas · Mid

Sponsorship not specifiedDetected 63 days ago
SQLSite Reliability EngineeringPlatform EngineeringIncident ResponseCadenceCommunication

About the role

  • How fast you detect them, how quickly you act, and whether the organization actually learns from them - that is what separates payments companies that scale from ones that spiral.
  • The question is not whether incidents will happen.
  • It is whether Forward detects them in minutes or hours, resolves them with coordination or chaos, and fixes the root cause or patches the symptom.

Responsibilities

  • Build the Detection Layer
  • Build and maintain AI-assisted signal intelligence: use AIOps platforms (PagerDuty, Incident.io, or equivalent) to correlate alerts, suppress noise, and surface high-confidence incident precursors before they manifest as partner escalations.
  • Own the governance loop over Support: review support ticket themes and escalation patterns on a weekly cadence to identify systemic issues before they cross into incident territory.
  • Establish alert-to-noise discipline: define what a true signal looks like for each incident type, tune alerting thresholds, and drive Alert-to-Noise Ratio above 80% - the team acts on signals, not volume.
  • Build and maintain the runbook library: pre-written, AI-augmented playbooks for the most common incident classes - submission failures, processing outages, ACH return spikes, TM system failures, compliance freezes - so the first 15 minutes of every incident are not spent figuring out who does what.
  • Serve as the named incident owner when an incident is declared - responsible for coordinating Engineering, Support, GTM, and Operations from detection through resolution.
  • Drive MTTD (Mean Time to Detect) and MTTA (Mean Time to Acknowledge) toward P1 targets: detection under 5 minutes, acknowledgment under 15 minutes for Sev-1 and Sev-2.
  • Classify merchant and partner impact in real time: GPV-at-risk, number of affected merchants, partner SLA implications, and any regulatory reporting obligations under DORA or card network rules.
  • A direct line to building the incident management function at a high-growth payments company from the ground up.
  • Per-partner MTTR trends down quarter-over-quarter.

Requirements

  • 4+ years in incident management, site reliability engineering, technical program management, or engineering operations - with direct ownership of production incident response.
  • Hands-on experience with modern incident management platforms: PagerDuty, Incident.io, FireHydrant, Rootly, Blameless, or equivalent AIOps tooling.
  • Familiarity with SLO/SLA frameworks, error budget concepts, and alert fatigue management.
  • internal stakeholder updates, partner-facing status, and merchant-level communications where required - proactive, not reactive.
  • Manage communications during incidents on a defined cadence: internal stakeholder updates, partner-facing status, and merchant-level communications where required - proactive, not reactive.

Nice to have

  • Familiarity with DORA regulatory incident reporting obligations or equivalent financial services operational resilience frameworks.
  • Experience integrating AI tooling into incident workflows: automated RCA, alert correlation, or runbook execution.
  • SQL proficiency for pulling operational data to support incident diagnosis and post-incident analysis.
  • Background in platform engineering or SRE at a payments or financial services company.
  • What Success Looks Like
  • Detection and Response
  • MTTD (Mean Time to Detect): under 5 minutes for Sev-1, under 15 minutes for Sev-2.
  • MTTA (Mean Time to Acknowledge): under 15 minutes for Sev-1, under 30 minutes for Sev-2.

Skills

  • Declare incidents using a consistent severity framework (Sev-1 through Sev-3) with defined, documented SLAs for each tier.
  • Drive

Compensation

  • Competitive salary and equity package.

Benefits

  • Competitive salary and equity package.
  • Comprehensive health, dental, and vision benefits.
  • Flexible work arrangements and generous PTO.
  • Learning & development budget for conferences, courses, and certifications.
  • SLO burn-rate alerting, synthetic health checks, deployment risk scoring, and real-time anomaly detection across submission, processing, and compliance pipelines.
  • monthly scorecards showing incident frequency, MTTR, and resolution quality by partner - inputs into GTM conversations and partner health reviews.

This listing is sourced directly from Forward's careers page and normalized into a canonical job model.