Gridware
Senior Cloud Engineer
San Francisco, California, USA · Senior · full-time
Sponsorship not specified$190k-$215kDetected 42 days ago
Distributed SystemsGitDatabricksAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringGraphQLKafkaMachine LearningData ScienceMLOpsCybersecurityIncident ResponseDNSZero Trust
About the role
- About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid.
- We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety.
- This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently.
Responsibilities
- For more information, please visit;br> Role Description We're scaling the deployment of critical infrastructure monitoring devices to detect real-world fault events that lead to wildfires.
- As an early member of the DevOps team, you'll have a direct hand in shaping how Gridware builds, deploys, and runs production systems for years to come.
- Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS.
- Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices.
- Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability.
- Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines.
- Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments.
- Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations.
- Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience.
- Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements.
Requirements
- Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production.
- Familiarity with Databricks or ML Ops pipelines for data and model deployment.
- Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks.
- Experience with Tailscale or other zero-trust networking tools.
- Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity.
- Experience in high-growth startup environments where you must wear many hats. $190,000 - $215,000 a year This describes the ideal candidate
- many of us have picked up this expertise along the way.
Compensation
- $190k-$215k
Benefits
- Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
- Commuter allowance Company-paid training
Company info
- Even if you meet only part of this list, we encourage you to apply!
Apply directly at Gridware →Create a free account for alerts like thisView Gridware immigration profile
This listing is sourced directly from Gridware's careers page and normalized into a canonical job model.