EarnIn

EarnIn

Site Reliability Engineer

Mountain View, US · Full-time

Sponsorship not specified$189k-$232kDetected 8 days ago
PythonGoDistributed SystemsGitDatadogSite Reliability EngineeringIncident ResponseLeadershipCommunicationMentoring

About the role

  • We're growing fast and are excited to continue bringing world-class talent onboard to help shape the next chapter of our growth journey.
  • Reliability shapes the product experience, not simply operational concerns.
  • Every noisy alert, unclear runbook, fragile deployment, or repeated incident undermines customer trust and hinders engineering teams.

Responsibilities

  • Design and improve systems with resilience and graceful degradation in mind. Plan for capacity and possible failure modes.
  • Use observability tools such as Datadog, CloudWatch, logs, metrics, traces, and APM. Build signal-heavy, noise-light visibility into production systems.
  • Construct or optimize infrastructure, reliability tooling, and automation that eliminate toil and ensure operational consistency.
  • Help engineering teams improve production readiness and deployment safety. Support service ownership and operational clarity.
  • Document operational knowledge to reduce silos. Make it easier for engineers to respond with confidence.
  • Design and improve systems with resilience and graceful degradation in mind.
  • You will meticulously validate AI-generated output before applying it to production systems or operational workflows.
  • You will collaborate with product engineering and platform teams to implement, explain, and support reliability practices, ensuring they are practical, understandable, and actionable.

Requirements

  • Bachelor's or master's degree in Computer Science, Engineering, or a related field, or equivalent industry experience.
  • 3+ years of experience in SRE, Software Engineering, Infrastructure Engineering, or a related role.
  • Hands-on coding experience in Python, Go, or similar production-oriented programming languages.
  • Experience operating production systems and contributing to reliability, observability, incident response, infrastructure, or automation improvements.
  • Working knowledge of SLIs, SLOs, error budgets, MTTR, and how reliability data informs engineering tradeoffs.
  • Experience using logs, metrics, dashboards, traces, and alerts to diagnose production issues.
  • Experience with distributed systems concepts such as retries, backoff, timeouts, graceful degradation, capacity planning, and failure isolation.
  • Experience improving alert quality, runbooks, incident processes, and follow-through after production issues.
  • Ability to communicate clearly, write useful documentation, and explain reliability concepts in plain language.
  • If you have any concerns about the use of AI in our hiring process, please connect with your recruiter.

Compensation

  • The base salary range for this full-time position is $189,000 - $232,000 plus equity and benefits.

Company info

  • WHAT we're looking for
  • At EarnIn, we believe that the best way to build a financial system that works for everyday people is by hiring a team that represents our diverse community.

Visa & Work Authorization

  • ildbirth, breastfeeding, or related medical conditions), gender identity, gender expression, national origin, ancestry, citizenship, age, physical or mental disability, legally protected medical condition, family care st

This listing is sourced directly from EarnIn's careers page and normalized into a canonical job model.