Synthesia

Synthesia

Senior Site Reliability Engineer

US Remote · Senior

Sponsorship not specifiedDetected 32 days ago
MongoDBAWSCloud PlatformsKubernetesDatadogSite Reliability EngineeringTemporalLLMsIncident ResponseValuationCommunication

About the role

  • We're hiring a dedicated SRE to take real ownership of operational excellence across Cloud Infrastructure.
  • Today, too much critical operational knowledge - vendor relationships, cost management, and incident response - lives with one or two people.
  • Your mission is to take genuine ownership of those domains, make them resilient to any single person, and raise the bar on how reliably we run.

Responsibilities

  • We're a small, high-leverage team scaling toward a domain-ownership model: small groups that both build and operate the systems they're accountable for.
  • You'll own domains end to end: understand them deeply, operate them well, and build the automation and tooling that make them boring.
  • Vendor & third-party management - own key external relationships and integrations (e.g. LLM API providers, third-party services), today managed manually and informally. Bring structure, automation, and bus-factor resilience.
  • You'll have a clear primary mission (operational excellence) but real domain ownership and the mandate to build - not a fixed lane.

Requirements

  • Strong production operations experience on AWS and Kubernetes

Nice to have

  • vendor/cost management exposure, Temporal, observability tooling.
  • How we think about this role
  • We don't letterbox engineers.
  • We expect the shape of the role to evolve as the team grows.
  • Synthesia is the world's leading AI video platform for business, used by over 90% of the Fortune 100.
  • Following our recent Series E funding round, where we raised $200 million, our valuation stands at $4 billion.

Skills

  • Critical operational knowledge is documented and shared - no single point of failure for vendor, cost, or incident response.

Company info

  • Cloud Infrastructure owns the platform every Synthesia product runs on - AWS, Kubernetes, MongoDB, Temporal, our observability stack, and the vendor and cost relationships underneath them.
  • High-risk manual processes are automated and self-documenting.
  • Measurable reliability gains: fewer SEV1-SEV3 incidents per quarter, faster customer-impact resolution, and a much higher share of incidents caught by monitoring before customers feel them.

This listing is sourced directly from Synthesia's careers page and normalized into a canonical job model.