Barbaricum

Barbaricum

Senior Reliability Engineer

Washington, DC · Senior · Contract

Sponsorship not specifiedDetected 34 days ago
PythonPowerShellAWSGCPAzureAnsibleLinuxSite Reliability EngineeringCybersecurityIncident ResponseSystems EngineeringLeadershipCommunicationCollaborationProblem Solving

About the role

  • Headquartered in Washington, DC's historic Dupont Circle neighborhood, Barbaricum also has a corporate presence in Tampa, FL, Bedford, IN, and Dayton, OH, with team members across the United States and around the world.
  • Through all of this, we have built a vibrant corporate culture diverse in expertise and perspectives with a focus on collaboration and innovation.

Responsibilities

  • Implement proactive performance monitoring, automated alerting, incident response workflows, and resilience engineering practices to reduce downtime and improve operational visibility.
  • Develop, maintain, and improve scalable automated infrastructure solutions that support reliable system operations and repeatable service delivery.
  • Implement rollback strategies, recovery approaches, and chaos engineering practices to validate resilience, reduce operational risk, and improve system stability.
  • Analyze usage patterns, capacity trends, and performance indicators to support dynamic scaling, resource optimization, and system improvement decisions.
  • Develop and maintain real-time operational dashboards, reports, and metrics that enable rapid decision-making, leadership awareness, and system optimization.
  • Conduct post-incident reviews to identify root causes, document lessons learned, and implement preventative measures that reduce recurrence.
  • Collaborate with software developers, cloud engineers, cybersecurity personnel, and operations teams to improve services, reliability patterns, deployment practices, and operational standards.
  • Create and maintain system documentation, configuration standards, operational runbooks, monitoring procedures, and service reliability guidance.
  • Implement security best practices across operational activities, infrastructure automation, monitoring, incident response, and system administration functions.

Requirements

  • Experience supporting incident response, system outage resolution, post-incident reviews, root cause analysis, and operational improvement initiatives.
  • Experience collaborating with development, infrastructure, cloud, cybersecurity, and program teams to improve reliability, security, and service performance.
  • Expert knowledge of site reliability engineering practices, system monitoring, incident management, automation, performance tuning, and operational resilience.
  • Strong understanding of Windows and Linux administration, infrastructure operations, system configuration, service management, and troubleshooting practices.
  • Proficiency with scripting languages such as Python, Shell, PowerShell, or similar tools used to automate operational and infrastructure tasks.
  • Strong understanding of network troubleshooting, configuration, connectivity analysis, system dependencies, and performance bottleneck identification.
  • Excellent written and verbal communication skills, with the ability to explain technical findings, incident impacts, and reliability recommendations to technical and non-technical stakeholders.

Nice to have

  • Master's degree preferred.
  • Bachelor's degree in Computer Science, Information Technology, Systems Engineering, Cybersecurity, or a related field

Skills

  • Experience with automation platforms and configuration management tools such as Ansible, Puppet, Chef, or similar technologies.
  • Knowledge of cloud services and infrastructure across AWS, Microsoft Azure, Google Cloud, or comparable cloud environments.
  • Ability to conduct root cause analysis, post-incident reviews, and corrective action planning in complex technical environments.
  • Strong problem-solving skills and the ability to work under pressure during outages, impairments, and time-sensitive operational issues.
  • Our teams are at the frontier of the Nation's most complex and rewarding challenges.

Visa & Work Authorization

  • DoD Secret Security Clearance

This listing is sourced directly from Barbaricum's careers page and normalized into a canonical job model.