Onebrief
Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided
United States | Remote · Senior
No sponsorshipDetected 14 days ago
PythonGoGitAWSKubernetesTerraformAnsibleCI/CDGitHub ActionsJenkinsPrometheusGrafanaDevOpsSite Reliability EngineeringCybersecurityIncident ResponseCollaborationAdaptability
About the role
- We are hiring a Site Reliability Engineer to join our Infrastructure & Security team.
- You'll work closely with fellow SREs, security, and customer success.
- You will work in both on-premise DoD environments and AWS cloud environments.
Responsibilities
- You'll own the reliability, scalability, and security of the production application and/or platform. You will do this by:
- Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics
- you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.
- Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally.
- Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.
- Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible).
- Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.
- Proven partner to DevOps/Platform and application teams
- collaborates well across functions and shares context openly.
- You'll own the reliability, scalability, and security of the production application and/or platform.
Requirements
- If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).
- Active Top Secret Clearance required with the ability to obtain SCI eligibility.
- proficiency with at least one of Python, Go, or Bash.
- Familiarity with AWS or AWS GovCloud.
Compensation
- In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee.
Company info
- An active Top Secret clearance
- 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
- Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.
- A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
- We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with fellow SREs, security, and customer success.
Apply directly at Onebrief →Create a free account for alerts like thisView Onebrief immigration profile
This listing is sourced directly from Onebrief's careers page and normalized into a canonical job model.