Onebrief

Onebrief

Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided

United States | Remote · Senior

No sponsorshipDetected 14 days ago
PythonGoGitAWSKubernetesTerraformAnsibleCI/CDGitHub ActionsJenkinsPrometheusGrafanaDevOpsSite Reliability EngineeringCybersecurityIncident ResponseCollaborationAdaptability

About the role

  • We are hiring a Site Reliability Engineer to join our Infrastructure & Security team.
  • You'll work closely with fellow SREs, security, and customer success.
  • You will work in both on-premise DoD environments and AWS cloud environments.

Responsibilities

  • You'll own the reliability, scalability, and security of the production application and/or platform. You will do this by:
  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics
  • you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.
  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally.
  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.
  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible).
  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.
  • Proven partner to DevOps/Platform and application teams
  • collaborates well across functions and shares context openly.
  • You'll own the reliability, scalability, and security of the production application and/or platform.

Requirements

  • If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).
  • Active Top Secret Clearance required with the ability to obtain SCI eligibility.
  • proficiency with at least one of Python, Go, or Bash.
  • Familiarity with AWS or AWS GovCloud.

Compensation

  • In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee.

Company info

  • An active Top Secret clearance
  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.
  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
  • We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with fellow SREs, security, and customer success.

This listing is sourced directly from Onebrief's careers page and normalized into a canonical job model.