Specter

Specter

Site Reliability Engineer

San Francisco

Sponsorship not specifiedDetected 19 days ago
PythonGoBashAWSCloud PlatformsDockerKubernetesLinuxSite Reliability EngineeringAgentic AIIncident ResponseEmbedded SystemsDNSFirewall

About the role

  • This is a high-ownership role at the intersection of ops and platform engineering.

Responsibilities

  • Own site bring-ups end to end
  • Identify repeat fires and eliminate them - build tooling, pre-deployment checks, and root cause processes that prevent recurrence.
  • Collaborate with embedded systems, and platform teams to define reliability and deployment requirements.
  • Design and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).
  • build fleet-wide visibility that enables data-driven reliability decisions.
  • Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments.
  • You'll drive reliability across our sensor fleet - triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale.

Requirements

  • Experience with edge or on-prem hardware alongside cloud infrastructure.
  • Strong Linux systems administration - comfortable working over SSH in production, not just dev environments.
  • Solid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.

Nice to have

  • Scripting or programming in Python, Go, or Bash for operational tooling.
  • Familiarity with containerization (Docker, Kubernetes a plus).
  • Embedded systems experience - reading firmware logs, understanding hardware-software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.
  • Deeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.
  • Rust or C experience - we have firmware in both
  • being able to read and reason about low-level code accelerates triage significantly.

Benefits

  • We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.
  • We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform - spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.
  • Build and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.

Company info

  • We offer both long range wireless (1km range) and wired sensor variants to suit any deployment.

This listing is sourced directly from Specter's careers page and normalized into a canonical job model.