Designworkstalent

Designworkstalent

Data Center Operations and Maintenance Engineering Leader

Bellevue

Sponsorship not specifiedDetected 19 hours ago
Cloud PlatformsPrometheusGrafanaDatadogAgentic AIIncident ResponseLeadershipCommunication

About the role

  • Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
  • Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging automation and modern tooling to operate world-class infrastructure at scale.
  • Depending on experience, responsibilities may range from leading a regional operations team to defining the long-term operational strategy across multiple facilities.

Responsibilities

  • Build, lead, mentor, and grow high-performing Operations & Maintenance teams.
  • Develop the operational strategy, organizational structure, and execution model supporting large-scale AI infrastructure.
  • Lead major incident management efforts and executive communications during production events.
  • Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement initiatives.
  • Partner with Engineering to ensure operational readiness for new infrastructure deployments and platform launches.
  • Build scalable on-call programs, escalation models, runbooks, and operational governance.
  • Proven success building or scaling operations teams within cloud infrastructure, hyperscale environments, AI infrastructure, or large distributed systems.
  • Build and lead the Operations & Maintenance organization from its earliest stages.

Requirements

  • Experience leading Operations, Site Reliability, Infrastructure Operations, Data Center Operations, or Production Engineering organizations.
  • Deep expertise in production operations, incident management, service reliability, and operational excellence.
  • Experience leading cross-functional teams during high-severity production incidents.
  • Strong understanding of infrastructure operations across compute, networking, storage, and hardware environments.
  • Executive-level communication skills with the ability to influence engineering and business leadership.
  • U.S. work authorization is required. Visa sponsorship is not currently available.

Nice to have

  • Experience supporting hyperscale cloud platforms, GPU infrastructure, AI platforms, HPC environments, or large-scale data centers.
  • Experience with modern observability and monitoring platforms such as Grafana, Prometheus, Datadog, or similar technologies.
  • Familiarity with incident management platforms including PagerDuty, Opsgenie, or equivalent solutions.
  • Experience implementing operational maturity frameworks and reliability engineering best practices.
  • Competitive base pay for Bellevue market
  • These awards are allocated based on individual performance
  • Approximately three days per week in the office.
  • U.S. work authorization is required.

Compensation

  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance

Benefits

  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Company info

  • As AI data centers come online across North America, the company is building its Operations & Maintenance leadership organization.

Visa & Work Authorization

  • U.S. work authorization is required.

This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.