Retool
Site Reliability Engineer (SRE)
San Francisco, USA
Sponsorship not specifiedDetected 14 days ago
TypeScriptPythonJavaGoPostgreSQLAWSDockerKubernetesTerraformHelmSite Reliability Engineering
About the role
- WHY WE'RE LOOKING FOR YOU: Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system. Retool's Core Infrastructure team owns the systems that make this
- possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers. The work is not clean-room infrastructure.
Responsibilities
- Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
- Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
- Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
Requirements
- Infrastructure fundamentals Deep experience operating production infrastructure in AWS.
- Experience improving reliability for customer-facing SaaS systems.
Skills
- The work is not clean-room infrastructure.
- A bad upgrade experience can leave a customer many versions behind.
- A manual Terraform run can become the bottleneck during a launch or incident.
- Some days that means debugging a specific customer environment.
- The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it.
Benefits
- We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
Company info
- Improve observability for Retool Cloud, self-hosted customers, and internal operators.
- Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
- Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.
- Good software has to run where customers need it.
- For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system.
This listing is sourced directly from Retool's careers page and normalized into a canonical job model.