Designworkstalent
Data Center Operations and Maintenance Engineering Leader
Bellevue
Sponsorship not specifiedDetected 19 hours ago
Cloud PlatformsPrometheusGrafanaDatadogAgentic AIIncident ResponseLeadershipCommunication
About the role
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
- Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging automation and modern tooling to operate world-class infrastructure at scale.
- Depending on experience, responsibilities may range from leading a regional operations team to defining the long-term operational strategy across multiple facilities.
Responsibilities
- Build, lead, mentor, and grow high-performing Operations & Maintenance teams.
- Develop the operational strategy, organizational structure, and execution model supporting large-scale AI infrastructure.
- Lead major incident management efforts and executive communications during production events.
- Drive operational excellence through proactive monitoring, observability, automation, and continuous improvement initiatives.
- Partner with Engineering to ensure operational readiness for new infrastructure deployments and platform launches.
- Build scalable on-call programs, escalation models, runbooks, and operational governance.
- Proven success building or scaling operations teams within cloud infrastructure, hyperscale environments, AI infrastructure, or large distributed systems.
- Build and lead the Operations & Maintenance organization from its earliest stages.
Requirements
- Experience leading Operations, Site Reliability, Infrastructure Operations, Data Center Operations, or Production Engineering organizations.
- Deep expertise in production operations, incident management, service reliability, and operational excellence.
- Experience leading cross-functional teams during high-severity production incidents.
- Strong understanding of infrastructure operations across compute, networking, storage, and hardware environments.
- Executive-level communication skills with the ability to influence engineering and business leadership.
- U.S. work authorization is required. Visa sponsorship is not currently available.
Nice to have
- Experience supporting hyperscale cloud platforms, GPU infrastructure, AI platforms, HPC environments, or large-scale data centers.
- Experience with modern observability and monitoring platforms such as Grafana, Prometheus, Datadog, or similar technologies.
- Familiarity with incident management platforms including PagerDuty, Opsgenie, or equivalent solutions.
- Experience implementing operational maturity frameworks and reliability engineering best practices.
- Competitive base pay for Bellevue market
- These awards are allocated based on individual performance
- Approximately three days per week in the office.
- U.S. work authorization is required.
Compensation
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
Benefits
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
Company info
- As AI data centers come online across North America, the company is building its Operations & Maintenance leadership organization.
Visa & Work Authorization
- U.S. work authorization is required.
Apply directly at Designworkstalent →Create a free account for alerts like thisView Designworkstalent immigration profile
This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.