Gridware

Gridware

Senior Data Analyst, Fleet Triage & Operations

San Francisco, CA · Senior

Sponsorship not specified$130k-$150kDetected 77 days ago
PythonDistributed SystemsSQLDatabricksGrafanaDatadogPandasData AnalysisData ScienceData VisualizationStatisticsDetection EngineeringIncident ResponseCommunication

About the role

  • As the deployed fleet scales, telemetry volume grows faster than any team can manually inspect, and not all signals matter equally.
  • This is a highly cross-functional role that blends data analysis, operational rigor, and structured incident learning across hardware, firmware, and connectivity domains.

Responsibilities

  • Maintain a tiered triage approach (P0-P3) with clear entry criteria, SLAs, ownership, and escalation paths across hardware, firmware, and connectivity stakeholders.
  • Build dashboards and queries that surface emerging trends - shifts in last-seen, connectivity degradation, configuration drift, anomalous failure clusters - before they reach incident severity.
  • Lead postmortems end-to-end: facilitate blameless reviews, reconstruct incident timelines, drive root cause analysis, and own action items through to closure.
  • Document triage playbooks, runbooks, and decision trees so the process scales beyond any one person.
  • This role separates signal from noise, ensures the right issues reach the right teams before they become incidents, and leads the postmortem process to make sure the same class of issue doesn't go undetected twice.

Requirements

  • Bachelor's degree in Computer Science, Data Science, Engineering, Statistics, or related field.
  • 5+ years in a data-focused or operations-focused role (Data Analyst, Operations Analyst, Site Reliability Analyst, or similar) where triage, prioritization, or incident response was a core responsibility.
  • Proficiency in SQL and Python for querying telemetry data and automating triage workflows.
  • Comfort facilitating cross-functional discussions, including blameless postmortems with engineering, firmware, and operations stakeholders.
  • Experience leading postmortems or incident reviews for IoT fleets, distributed systems, or large-scale hardware deployments.
  • Familiarity with Databricks, PySpark, Pandas, and observability tools (Looker, Grafana, Datadog).
  • Experience with power-constrained, solar-powered, or communication-constrained connected devices.
  • This describes the ideal candidate
  • many of us have picked up this expertise along the way.

Compensation

  • Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.

Benefits

  • Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
  • Commuter allowance Company-paid training
  • Own and operate Gridware's fleet triage pipeline for device health issues, intaking signals from telemetry, monitoring, field reports, and customer feedback, and routing them through a tiered prioritization framework.
  • Define and continuously tune detection logic, thresholds, and alerting rules that distinguish signal from noise across device health, connectivity, configuration, and environmental data.

This listing is sourced directly from Gridware's careers page and normalized into a canonical job model.