Certifyos
Senior Site Reliability Engineer
Remote US · Senior
Sponsorship not specifiedDetected 34 days ago
TypeScriptPythonJavaBashReactNode.jsDistributed SystemsGitBigQueryGCPCloud PlatformsDockerKubernetesTerraformCI/CDGitHub ActionsLinuxPrometheusGrafanaDatadogDevOpsSite Reliability EngineeringPlatform EngineeringIncident Response
About the role
- You'll work across cloud-native infrastructure on systems that process millions of provider records.
- We use AI-assisted tooling aggressively to reduce toil and accelerate troubleshooting, which raises the floor on the problems we tackle - not an excuse to reduce rigor.
- If you do your best work reacting to incidents, this probably isn't the right fit.
Responsibilities
- This creates unnecessary administrative work, regulatory risk, and higher costs across the system.
- This is a role with real scope: you'll own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows across systems that matter.
- SREs at CertifyOS own the full lifecycle of what they support - from infrastructure design and deployment automation through observability, incident response, and postmortems.
- You'll own incident response processes, root cause analysis, escalation workflows, and runbooks - and make hard problems not happen again.
- You'll build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling that reduce manual work and improve engineering productivity without sacrificing reliability.
- At Certify, we're building with intention and taking care of the people doing the work.
- We're building a high-ownership team focused on solving real infrastructure problems that impact millions of patients.
- you'll own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows across systems that matter.
Requirements
- 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering - operating production systems at scale where your infrastructure is someone else's dependency and failures have real downstream consequences
- Track record of improving reliability end-to-end: you've debugged hard production problems, made them not happen again, and built the alerting to prove it
- Comfort influencing operational standards and mentoring teams on reliability practices
- Deep hands-on experience with GCP - GKE, Cloud Run, and containerized workloads at scale
- Experience with autoscaling, resource optimization, and infrastructure efficiency for distributed systems
- Experience managing infrastructure security, secrets, and access controls in regulated or security-conscious environments
- Strong understanding of Golden Signals monitoring - latency, traffic, errors, saturation - and how to make them actionable rather than noisy
- Experience designing SLIs, SLOs, error budgets, alerting strategies, dashboards, and escalation workflows
- Hands-on experience with observability platforms: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar
- Experience working with Git workflows and modern software delivery practices
Nice to have
- Experience operating large-scale distributed systems or microservices architectures
- Experience leveraging AI-assisted observability or incident response tooling
- Familiarity with NodeJS, TypeScript, Java, or React application stacks
- Your well-being matters to us.
Compensation
- We are also committed to pay transparency and foster an open culture where compensation conversations are encouraged and respected.
Benefits
- We provide 100% coverage of health, dental, and vision insurance premiums for employees.
- Our US-based team benefits from unlimited PTO, with at least two weeks off each year to recharge.
- In India, employees are supported with health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women.
- CertifyOS is building the data infrastructure that powers modern healthcare.
- Today, healthcare organizations rely on fragmented and outdated provider data.
- We help healthcare organizations maintain accurate, compliant, and reliable provider networks at scale.
Company info
- Our API-first platform automates provider licensing, enrollment, credentialing, and network monitoring by connecting directly to hundreds of primary data sources.
- We are also committed to pay transparency and foster an open culture where compensation conversations are encouraged and respected.
- At CertifyOS, we value authenticity, accountability, collaboration, results, and openness to feedback.
Equal opportunity
- We are an equal opportunity employer committed to building an inclusive environment where everyone feels valued and empowered to do their best work, and we welcome applicants from all backgrounds and experiences.
Apply directly at Certifyos →Create a free account for alerts like thisView Certifyos immigration profile
This listing is sourced directly from Certifyos's careers page and normalized into a canonical job model.