Synthesia
Senior Site Reliability Engineer
US Remote · Senior
Sponsorship not specifiedDetected 32 days ago
MongoDBAWSCloud PlatformsKubernetesDatadogSite Reliability EngineeringTemporalLLMsIncident ResponseValuationCommunication
About the role
- We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure.
- Today, too much critical operational knowledge - vendor relationships, cost management, and incident response - lives with one or two people.
- Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run.
Responsibilities
- We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for.
- You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring.
- Vendor & third-party management - own key external relationships and integrations (e.g. LLM API providers, third-party services), today managed manually and informally. Bring structure, automation, and bus-factor resilience.
- You'll have a clear primary mission (operational excellence) but real domain ownership and the mandate to build - not a fixed lane.
Requirements
- Strong production operations experience on AWS and Kubernetes
Nice to have
- vendor/cost management exposure, Temporal, observability tooling.
- How we think about this role
- We don't letterbox engineers.
- We expect the shape of the role to evolve as the team grows.
- Synthesia is the world's leading AI video platform for business, used by over 90% of the Fortune 100.
- Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion.
Skills
- Critical operational knowledge is documented and shared - no single point of failure for vendor, cost, or incident response.
Company info
- Cloud Infrastructure owns the platform every Synthesia product runs on - AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them.
- High-risk manual processes are automated and self-documenting.
- Measurable reliability gains: fewer SEV1-SEV3 incidents per quarter, faster customer-impact resolution, and a much higher share of incidents caught by monitoring before customers feel them.
Apply directly at Synthesia →Create a free account for alerts like thisView Synthesia immigration profile
This listing is sourced directly from Synthesia's careers page and normalized into a canonical job model.