Gridware

Gridware

Senior Cloud Engineer

San Francisco, CA · Senior

Sponsorship not specified$190k-$215kDetected 76 days ago
Distributed SystemsGitDatabricksAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringGraphQLKafkaMachine LearningData ScienceMLOpsCybersecurityIncident ResponseDNSZero Trust

About the role

  • We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety.

Responsibilities

  • Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS.
  • Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices.
  • Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability.
  • Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines.
  • Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments.
  • Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations.
  • Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience.
  • Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements.
  • Strong experience building and maintaining CI/CD pipelines, ideally with GitHub Actions

Requirements

  • 5+ years of experience in DevOps, SRE, or Platform Engineering operating production AWS environments
  • Hands-on experience operating distributed systems and cloud-native platforms (e.g., Kafka/MSK)
  • Experience with observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar
  • Strong Linux, scripting, and troubleshooting skills with the ability to debug complex production issues end-to-end
  • Experience operating Apollo Router / GraphQL federation gateways in production.
  • Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production.
  • Familiarity with Databricks or ML Ops pipelines for data and model deployment.
  • Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks.
  • Experience with Tailscale or other zero-trust networking tools.
  • Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity.

Nice to have

  • Deep expertise with Kubernetes (EKS preferred), GitOps workflows (Argo CD/Flux), and Infrastructure as Code (Terraform)

Skills

  • About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid.
  • This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently.
  • The company is backed by climate-tech and Silicon Valley investors.
  • For more information, please visit www.Gridware.io.

Compensation

  • Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.

Benefits

  • Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
  • Commuter allowance Company-paid training

This listing is sourced directly from Gridware's careers page and normalized into a canonical job model.