Veeamsoftware

Senior Site Reliability Engineer- FedRamp

Remote, United States · Senior

Sponsorship not specifiedDetected 1 day ago
JavaScriptTypeScriptJavaGoC#Distributed SystemsGitAWSAzureCloud PlatformsKubernetesTerraformCI/CDGitHub ActionsPrometheusGrafanaDevOpsSite Reliability EngineeringPlatform EngineeringLLMsCybersecurityIncident ResponseComplianceHIPAA

About the role

  • This role focuses on our Government and Sovereign Cloud environment.
  • That means you'll be part of a small team responsible for the full platform stack - including all VDC workloads.
  • You won't always be able to hand off problems to other teams

Responsibilities

  • Work with SMEs across the org to fill knowledge gaps and build onboarding material for the team.
  • Write and maintain runbooks, architecture docs, and operational guides.
  • Design infrastructure for high availability and fault tolerance on Azure (including Azure Government).
  • Identify reliability risks across modern and legacy workloads and build practical remediation plans that work within compliance constraints.
  • Close observability gaps - define instrumentation requirements and drive implementation.
  • Set alerting, telemetry, and monitoring standards with partner teams.
  • Build automation to reduce toil and support fleet management.

Requirements

  • Experience with Government or Sovereign Cloud (e.g., Azure Government, AWS GovCloud).

Skills

  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK stack).
  • Experience with IaC (Terraform, Terragrunt, Pulumi) and container orchestration (Kubernetes).
  • Experience with CI/CD and GitOps tooling - GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger.
  • Solid grasp of distributed systems, networking, and cloud-native architecture.
  • Clear written and verbal communication skills
  • Experience on B2B SaaS platforms in regulated or government markets.
  • Background in chaos engineering, resilience testing, or performance/load testing.
  • Have built an SRE or reliability function from scratch before.
  • Experience across mixed environments - modern cloud-native and older legacy systems.
  • Familiar with AI-first development workflows - using LLM-powered tools for infrastructure automation, code generation, and documentation.

Company info

  • Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale.
  • As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk.
  • Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running.
  • Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world's biggest brands.

This listing is sourced directly from Veeamsoftware's careers page and normalized into a canonical job model.