Gridware
Senior Cloud Engineer
San Francisco, CA · Senior
Sponsorship not specified$190k-$215kDetected 76 days ago
Distributed SystemsGitDatabricksAWSCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsLinuxPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringGraphQLKafkaMachine LearningData ScienceMLOpsCybersecurityIncident ResponseDNSZero Trust
About the role
- We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety.
Responsibilities
- Design, build, and operate scalable, secure, and highly available cloud infrastructure across AWS.
- Own and evolve our Kubernetes platform, enabling reliable application deployment and operations through GitOps best practices.
- Build and maintain CI/CD systems that improve developer velocity, release quality, and operational reliability.
- Manage and optimize event-driven infrastructure powering high-volume telemetry and device data pipelines.
- Define and maintain Infrastructure as Code standards, ensuring consistency, repeatability, and scalability across environments.
- Develop and enhance observability, monitoring, and incident response capabilities to support reliable production operations.
- Partner closely with Security and Engineering teams to strengthen platform security, access management, and operational resilience.
- Troubleshoot complex production issues, drive root cause analysis, and turn lessons learned into automation, tooling, and operational improvements.
- Strong experience building and maintaining CI/CD pipelines, ideally with GitHub Actions
Requirements
- 5+ years of experience in DevOps, SRE, or Platform Engineering operating production AWS environments
- Hands-on experience operating distributed systems and cloud-native platforms (e.g., Kafka/MSK)
- Experience with observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar
- Strong Linux, scripting, and troubleshooting skills with the ability to debug complex production issues end-to-end
- Experience operating Apollo Router / GraphQL federation gateways in production.
- Experience operating Argo Workflows or similar Kubernetes-native job / pipeline runners in production.
- Familiarity with Databricks or ML Ops pipelines for data and model deployment.
- Experience designing, operating, and exercising Disaster Recovery (DR) environments, including cross-region replication, backups, and tested failover runbooks.
- Experience with Tailscale or other zero-trust networking tools.
- Experience supporting IoT / embedded fleets at scale, including secure device-to-cloud connectivity.
Nice to have
- Deep expertise with Kubernetes (EKS preferred), GitOps workflows (Argo CD/Flux), and Infrastructure as Code (Terraform)
Skills
- About Gridware Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid.
- This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently.
- The company is backed by climate-tech and Silicon Valley investors.
- For more information, please visit www.Gridware.io.
Compensation
- Benefits Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
Benefits
- Health, Dental & Vision (Gold and Platinum with some providers plans fully covered) Paid parental leave Alternating day off (every other Monday) "Off the Grid", a two week per year paid break for all employees.
- Commuter allowance Company-paid training
Apply directly at Gridware →Create a free account for alerts like thisView Gridware immigration profile
This listing is sourced directly from Gridware's careers page and normalized into a canonical job model.