ServiceNow
Staff Reliability Engineer
Santa Clara, CALIFORNIA, United States · Staff+ · Contract
Sponsorship not specified$167k-$291kDetected 5 hours ago
PythonJavaGoDistributed SystemsCode ReviewGitAWSGCPAzureKubernetesTerraformAnsibleHelmCI/CDPrometheusDevOpsSite Reliability EngineeringPlatform EngineeringRESTSeleniumCypressPlaywrightpytestJUnit
> stay_score
odds of building a lasting career here
37Unrated
Cap-exempt (no lottery)0
Sponsors this role0
Entry-level history0
PERM / green-card track0
Lottery odds (Level IV)94
Fits your clock70
No strong sponsorship signal in the public record yet. In the full product we resolve the exact legal entity and show its filing history with a confidence score — treat as unverified until then.
Lottery odds assume a STEM candidate.
Personalize to your clock →> community_outcomes
No reports yet — be the first to help the next applicant.
About the role
- It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis.
- To be successful in this role you have: - Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
Responsibilities
- Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness
- Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness.
- Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows.
- Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices.
- Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling.
- Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service.
- Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation.
- Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices.
- Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions.
- Experience building and operating cloud-native platforms supporting scalable, highly available services.
Requirements
- Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments.
- Experience with progressive delivery practices, including canary deployments, feature flags, automated rollback, and deployment verification.
- Experience with chaos engineering, resilience testing, disaster recovery, and reliability validation.
- Strong understanding of observability, monitoring, SLI/SLOs, incident management, and production operations for distributed systems.
- To be successful in this role you have:.
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
- 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree
- or 6 years and a Master's degree
- or a PhD with 3 years experience
Nice to have
- Experience with observability and monitoring platforms for applications, services, and distributed systems at scale.
- Experience with DevOps automation, CI/CD pipelines, GitOps, and Agile development practices using tools such as GitLab CI/CD, Argo CD, or Flux.
- Experience with test orchestration, intelligent regression testing, test impact analysis, flaky test detection, parallel execution, and test data management.
- Experience with service virtualization, contract testing, synthetic testing, and building developer self-service engineering platforms.
- Experience with Infrastructure as Code and configuration management tools such as Ansible, Terraform, or equivalent.
- Experience with the Kubernetes ecosystem, including Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtime technologies.
- Experience operating Kubernetes platforms across public cloud providers, including AWS (EKS), Azure (AKS), and Google Cloud (GKE).
- Experience implementing progressive delivery practices, including canary deployments, feature flags, deployment verification, and automated rollback.
Skills
- She cried tears of joy.
- Today, ServiceNow is the AI control tower for business reinvention.
- Join us to put AI to work for people.
Compensation
- $167k-$291k
Benefits
- Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments.
- Thrives in fast-paced, ambiguous environments with a strong ownership mindset, bias for action, and a passion for continuous learning and automation.
Equal opportunity
- ServiceNow is an equal opportunity employer.
- All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law.
- In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
- Accommodations
Visa & Work Authorization
- Export Control Regulations
Apply directly at ServiceNow →Create a free account for alerts like thisView ServiceNow immigration profile
This listing is sourced directly from ServiceNow's careers page and normalized into a canonical job model.