Gridware

Gridware

Senior Cloud Engineer

San Francisco, California, USA · Senior · full-time

Sponsorship not specified$190k-$215kDetected 42 days ago
Distributed SystemsGitDatabricksAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringGraphQLKafkaMachine LearningData ScienceMLOpsCybersecurityIncident ResponseDNSZero Trust

About the role

  • About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid.
  • We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety.
  • This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently.

Responsibilities

  • For more information, please visit;br> Role Description We're scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires.
  • As an early member of the DevOps team, you'll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come.
  • Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS.
  • Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices.
  • Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability.
  • Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines.
  • Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments.
  • Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations.
  • Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience.
  • Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements.

Requirements

  • Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production.
  • Familiarity with Databricks or ML Ops pipelines for data and model deployment.
  • Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks.
  • Experience with Tailscale or other zero-trust networking tools.
  • Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity.
  • Experience in high-growth startup environments where you must wear many hats. $190,000 - $215,000 a year This describes the ideal candidate
  • many of us have picked up this expertise along the way.

Compensation

  • $190k-$215k

Benefits

  • Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
  • Commuter allowance Company-paid training

Company info

  • Even if you meet only part of this list, we encourage you to apply!

This listing is sourced directly from Gridware's careers page and normalized into a canonical job model.