Specter
Site Reliability Engineer
San Francisco
Sponsorship not specifiedDetected 19 days ago
PythonGoBashAWSCloud PlatformsDockerKubernetesLinuxSite Reliability EngineeringAgentic AIIncident ResponseEmbedded SystemsDNSFirewall
About the role
- This is a high-ownership role at the intersection of ops and platform engineering.
Responsibilities
- Own site bring-ups end to end
- Identify repeat fires and eliminate them - build tooling, pre-deployment checks, and root cause processes that prevent recurrence.
- Collaborate with embedded systems, and platform teams to define reliability and deployment requirements.
- Design and implement observability (logging, metrics, alerting) across edge devices and cloud infrastructure (AWS).
- build fleet-wide visibility that enables data-driven reliability decisions.
- Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments.
- You'll drive reliability across our sensor fleet - triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale.
Requirements
- Experience with edge or on-prem hardware alongside cloud infrastructure.
- Strong Linux systems administration - comfortable working over SSH in production, not just dev environments.
- Solid networking fundamentals: DNS, firewalls, VPNs, subnets, secure remote access.
Nice to have
- Scripting or programming in Python, Go, or Bash for operational tooling.
- Familiarity with containerization (Docker, Kubernetes a plus).
- Embedded systems experience - reading firmware logs, understanding hardware-software boundaries, and reasoning about what's happening below the OS is a meaningful edge in this role.
- Deeper cloud experience (AWS infrastructure, IAM, networking, observability tooling) is a strong plus for owning the cloud side of the fleet.
- Rust or C experience - we have firmware in both
- being able to read and reason about low-level code accelerates triage significantly.
Benefits
- We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.
- We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform - spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.
- Build and maintain fleet management systems: OTA update pipelines, device health tracking, remote diagnostics, and lifecycle tooling.
Company info
- We offer both long range wireless (1km range) and wired sensor variants to suit any deployment.
Apply directly at Specter →Create a free account for alerts like thisView Specter immigration profile
This listing is sourced directly from Specter's careers page and normalized into a canonical job model.