PostHog

PostHog

Site Reliability Engineer (US - Central/Eastern time)

Remote (US)

Sponsorship not specifiedDetected 6 days ago
Node.jsGitSQLAWSKubernetesTerraformCI/CDGitHub ActionsLinuxSite Reliability EngineeringAgentic AIIncident ResponseCommunicationArgo CD

About the role

  • ABOUT POSTHOG Product development used to mean manually writing code, running analysis, diagnosing bugs, and rolling out changes using dozens of tools.
  • We started with open-source product analytics, launched out of Y Combinator's W20 cohort https://posthog.com/handbook/story.
  • More than 450,000 organizations have installed PostHog, mostly driven by word-of-mouth.

Responsibilities

  • That means reducing operational stress, designing safe automation for traffic-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort.
  • You'll have room to design and automate, not just respond to alerts.
  • You should join this team if you like deep ownership of production systems and enjoy building the platform layer that everything else runs on.
  • Why not now? https://posthog.com/handbook/values#why-not-now We want to build a lot of products; we can't do that shipping at a normal pace.
  • We've built the company around small teams - autonomous, highly-efficient groups of cracked engineers https://posthog.com/founders/cracked-manifesto who can outship much larger companies because they own their products end-to-end.
  • Enthusiastic drivers. We need proactive people that can fully own projects and get them done, and know to get help when needed. "Are we there yet?" is the wrong question.
  • Things get hard here sometimes, whether it's scaling, shipping complex products, handling a stream of support requests, or trying to ship something that touches multiple teams.
  • We need people who won't get disheartened, and will collaborate, iterate, and ship their way out of anything.

Requirements

  • Strong experience operating production infrastructure on AWS.
  • Experience supporting stateful systems (databases, queues, storage systems, etc.)
  • Ability to debug and reason about performance and reliability issues in production
  • We're optimistic about what's possible and our ability to get there.

Nice to have

  • Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
  • Experience with building AI agent-enabled base-level infra services for teams that move fast
  • Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it
  • If you need any accommodations or adjustments, please let us know.
  • Deep hands-on experience with Kubernetes in production (EKS preferred).
  • You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)

Skills

  • PostHog Code https://posthog.com/code, the only AI devtool that understands your product, not just your codebase.
  • Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
  • Managing and evolving a multi AWS account organization, provisioning, networking, access control, and cross-account connectivity
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
  • Optimizing cloud spend as you go
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time

Benefits

  • The work is about turning a fast-growing, stateful system into a predictable, well-automated platform. (provisioning, scaling, rebalancing, recovery)

Company info

  • Product-led. More than 450,000 organizations have installed PostHog, mostly driven by word-of-mouth. We have intensely strong product-market fit.
  • Default alive https://paulgraham.com/aord.html. Revenue is growing incredibly quickly, and we're very efficient. We raise money to push ambition and grow faster, not to keep the lights on.
  • Well-funded. We've raised more than $180m from some of the world's top investors. We're set up for a long, ambitious journey.
  • Everyone can read about our roadmap, how we pay (or even let go of) people, our strategy, and how we work, in our public company handbook https://posthog.com/handbook.
  • Internally, we share revenue, notes and slides from board meetings, and fundraising plans, so everyone has the context they need to make good decisions.
  • We don't tell anyone what to do.
  • Everyone chooses what to work on next based on what's going to have the biggest impact on our customers, and what they find interesting and motivating to work on.
  • Engineers lead product teams https://posthog.com/handbook/wide-company and make product decisions https://posthog.com/handbook/which-products.
  • Teams are flexible and easy to change when needed.
  • Nothing gets shipped in a meeting.
  • We're a natively remote company.
  • We default to async communication - PRs > Issues > Slack.
  • Tuesdays and Thursdays are meeting-free days https://posthog.com/handbook/company/culture#were-on-the-makers-schedule, and we prioritize heads down building time over perfect coordination.
  • This will be the most productive job you've ever had.

This listing is sourced directly from PostHog's careers page and normalized into a canonical job model.