OpsMill

OpsMill

Product Reliability Engineer

US/Canada

Sponsorship not specifiedDetected 173 days ago
PythonGoDistributed SystemsCode ReviewKubernetesHelmCI/CDSite Reliability EngineeringPlatform EngineeringCustomer SuccessTest AutomationCommunication

About the role

  • Shipping infrastructure software is only half the job.
  • The difference between a good product and a trusted one is how quickly you can diagnose issues and how effectively you prevent them from happening again.
  • WHY THIS ROLE EXISTS We need someone who can operate in both worlds: diving deep on gnarly customer escalations while systematically eliminating entire classes of problems.

Responsibilities

  • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering-then documenting learnings in crisp RCAs that become actionable improvements
  • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early
  • When something breaks in the field, it's not just a support ticket-it's a signal about what we need to fix, test, or instrument better.
  • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do
  • Strong software engineering fundamentals including design, debugging, testing, code review, and a focus on maintainable, production-quality code

Requirements

  • 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you've owned reliability and customer escalations
  • Practical Kubernetes expertise sufficient to debug real deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific quirks across different distributions
  • Deep troubleshooting instincts and observability experience using logs, metrics, and traces to diagnose issues quickly in complex, distributed systems
  • Experience with at least one of: Python, Go, or Rust for building tooling and contributing to product code (you don't need to be expert in all three)
  • Excellent problem decomposition and communication skills-you can break down messy, ambiguous issues and clearly explain your findings and recommendations
  • Experience with packaging and distribution systems (containers, Helm charts, installers) and managing upgrade/migration flows
  • Familiarity with performance tooling such as profiling, load generation, and benchmark harnesses

Benefits

  • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster

Company info

  • Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations-communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments.
  • At OpsMill, we're building Infrahub, a schema-driven infrastructure source of truth that helps teams unify data and scale automation reliably.
  • Our customers deploy Infrahub on-prem, which means reliability is a product feature, not just an operational concern.

This listing is sourced directly from OpsMill's careers page and normalized into a canonical job model.