Specter
Fleet Reliability Engineer
San Francisco
Sponsorship not specifiedDetected 8 days ago
PythonSQLPostgreSQLTerraformPrometheusGrafanaDatadogSite Reliability EngineeringAgentic AIForecasting
About the role
- As we scale toward thousands of sensors, fleet health becomes a data-and-systems problem.
- Much of today's operational load is addressable through better instrumentation, alert hygiene, and recovery verification, at little to no field cost.
Responsibilities
- Own the fleet's reliability data pipeline end to end: telemetry aggregation, storage, and instrumentation.
- Drive down observability cost - own the tooling spend and cut what we pay for but don't use.
- Build the failure-mode analysis that tells engineering what to fix at the source.
- Own fleet-wide trend and forecasting work, including power and solar planning.
- Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments.
- We're hiring a Fleet Reliability Engineer to keep our sensor fleet running in the field by building the data, analytics, and recovery mechanisms that prevent failures from becoming incidents.
Requirements
- Experience operating physical or embedded device fleets at scale, and reasoning about how hardware fails in the field.
- Strong data and software skills - Python (or Go) and SQL - and the ability to own a data pipeline end to end.
- Hands-on building and tuning observability stacks (OpenTelemetry, Grafana, Prometheus, Datadog, or similar), including their cost.
- Comfortable turning messy field telemetry into trends, failure modes, and forecasts.
- Bias toward building mechanisms over doing manual work.
Nice to have
- reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.
- experience across the hardware-software boundary - power, connectivity, and physical failure modes.
- Fluency with databases and data modeling (PostgreSQL or equivalent)
- infrastructure-as-code familiarity (Terraform or similar) a plus.
- Nice to have: reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.
- Nice to have: experience across the hardware-software boundary - power, connectivity, and physical failure modes.
Benefits
- We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.
- Instrument the fleet and own the health metrics that measure reliability.
- Fleet Health & Failure-Mode Analytics
Company info
- We offer both long range wireless (1km range) and wired sensor variants to suit any deployment.
Apply directly at Specter →Create a free account for alerts like thisView Specter immigration profile
This listing is sourced directly from Specter's careers page and normalized into a canonical job model.