OpsMill
Product Reliability Engineer
US/Canada
Sponsorship not specifiedDetected 173 days ago
PythonGoDistributed SystemsCode ReviewKubernetesHelmCI/CDSite Reliability EngineeringPlatform EngineeringCustomer SuccessTest AutomationCommunication
About the role
- Shipping infrastructure software is only half the job.
- The difference between a good product and a trusted one is how quickly you can diagnose issues and how effectively you prevent them from happening again.
- WHY THIS ROLE EXISTS We need someone who can operate in both worlds: diving deep on gnarly customer escalations while systematically eliminating entire classes of problems.
Responsibilities
- Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering-then documenting learnings in crisp RCAs that become actionable improvements
- Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early
- When something breaks in the field, it's not just a support ticket-it's a signal about what we need to fix, test, or instrument better.
- Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do
- Strong software engineering fundamentals including design, debugging, testing, code review, and a focus on maintainable, production-quality code
Requirements
- 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you've owned reliability and customer escalations
- Practical Kubernetes expertise sufficient to debug real deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific quirks across different distributions
- Deep troubleshooting instincts and observability experience using logs, metrics, and traces to diagnose issues quickly in complex, distributed systems
- Experience with at least one of: Python, Go, or Rust for building tooling and contributing to product code (you don't need to be expert in all three)
- Excellent problem decomposition and communication skills-you can break down messy, ambiguous issues and clearly explain your findings and recommendations
- Experience with packaging and distribution systems (containers, Helm charts, installers) and managing upgrade/migration flows
- Familiarity with performance tooling such as profiling, load generation, and benchmark harnesses
Benefits
- Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster
Company info
- Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations-communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments.
- At OpsMill, we're building Infrahub, a schema-driven infrastructure source of truth that helps teams unify data and scale automation reliably.
- Our customers deploy Infrahub on-prem, which means reliability is a product feature, not just an operational concern.
Apply directly at OpsMill →Create a free account for alerts like thisView OpsMill immigration profile
This listing is sourced directly from OpsMill's careers page and normalized into a canonical job model.