Specter

Specter

Fleet Reliability Engineer

San Francisco

Sponsorship not specifiedDetected 8 days ago
PythonSQLPostgreSQLTerraformPrometheusGrafanaDatadogSite Reliability EngineeringAgentic AIForecasting

About the role

  • As we scale toward thousands of sensors, fleet health becomes a data-and-systems problem.
  • Much of today's operational load is addressable through better instrumentation, alert hygiene, and recovery verification, at little to no field cost.

Responsibilities

  • Own the fleet's reliability data pipeline end to end: telemetry aggregation, storage, and instrumentation.
  • Drive down observability cost - own the tooling spend and cut what we pay for but don't use.
  • Build the failure-mode analysis that tells engineering what to fix at the source.
  • Own fleet-wide trend and forecasting work, including power and solar planning.
  • Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments.
  • We're hiring a Fleet Reliability Engineer to keep our sensor fleet running in the field by building the data, analytics, and recovery mechanisms that prevent failures from becoming incidents.

Requirements

  • Experience operating physical or embedded device fleets at scale, and reasoning about how hardware fails in the field.
  • Strong data and software skills - Python (or Go) and SQL - and the ability to own a data pipeline end to end.
  • Hands-on building and tuning observability stacks (OpenTelemetry, Grafana, Prometheus, Datadog, or similar), including their cost.
  • Comfortable turning messy field telemetry into trends, failure modes, and forecasts.
  • Bias toward building mechanisms over doing manual work.

Nice to have

  • reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.
  • experience across the hardware-software boundary - power, connectivity, and physical failure modes.
  • Fluency with databases and data modeling (PostgreSQL or equivalent)
  • infrastructure-as-code familiarity (Terraform or similar) a plus.
  • Nice to have: reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.
  • Nice to have: experience across the hardware-software boundary - power, connectivity, and physical failure modes.

Benefits

  • We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.
  • Instrument the fleet and own the health metrics that measure reliability.
  • Fleet Health & Failure-Mode Analytics

Company info

  • We offer both long range wireless (1km range) and wired sensor variants to suit any deployment.

This listing is sourced directly from Specter's careers page and normalized into a canonical job model.