OpenAI

OpenAI

Hardware Technical Program Manager, Infrastructure Partner Operations

San Francisco · Contract

Sponsorship not specifiedDetected 2 days ago
AWSGCPAzureCloud PlatformsIncident ResponseAccount ManagementSupply ChainProcess ImprovementCustomer SuccessHardware DesignResearchCommunicationPublic Speaking

About the role

  • As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical.

Responsibilities

  • Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations.
  • Develop operational governance frameworks with strategic partners, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
  • Build dashboards and reporting mechanisms that provide clear visibility into partner operational performance, risks, trends, and areas requiring executive attention.
  • Drive cross-functional coordination between OpenAI teams and external infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
  • Lead operational escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication.
  • Establish repeatable operating rhythms with external partners, including weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives.
  • Partner with Capacity Planning, Hardware Operations, Networking, Deployment, Reliability Engineering, and Supply Chain teams to ensure external infrastructure providers remain aligned with OpenAI's operational priorities.
  • Identify systemic operational risks across partner organizations and proactively drive corrective actions that improve long-term operational effectiveness.

Requirements

  • 7+ years of experience in Technical Program Management, Infrastructure Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments.
  • Strong understanding of hyperscale cloud infrastructure, data center operations, infrastructure delivery, or large-scale distributed systems.
  • Bachelor's degree in Engineering, Computer Science, Information Systems, Operations, or equivalent practical experience.
  • Strong understanding of service-level agreements (SLAs), operational KPIs, incident management, escalation processes, root cause analysis, and continuous service improvement methodologies.
  • Familiarity with infrastructure operations supporting GPU infrastructure, AI infrastructure, high-performance computing (HPC), or hyperscale data center environments.
  • Experience managing operational relationships with external infrastructure providers, cloud service providers, hardware vendors, or strategic technology partners.
  • Experience developing operational KPIs, SLAs, service health metrics, dashboards, and executive reporting for complex technical organizations.
  • Demonstrated success leading cross-functional operational programs involving both internal stakeholders and external partners.
  • Strong program management skills with the ability to drive accountability across organizations without direct authority.
  • Excellent written and verbal communication skills with experience presenting operational performance to senior technical and executive leadership.
  • Experience managing cloud infrastructure operations within organizations such as Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), or other hyperscale cloud providers.
  • Experience leading operational governance, service delivery, customer success engineering, technical account management, or infrastructure operations for enterprise cloud customers.
  • Experience building executive dashboards, operational scorecards, business review frameworks, and data-driven performance reporting.

Nice to have

  • Preferred Skills

Benefits

  • Define, track, and continuously improve key operational metrics related to infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.

Company info

  • The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI's largest-scale AI systems.
  • We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads.
  • Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI's rapidly expanding compute environment.
  • We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.

This listing is sourced directly from OpenAI's careers page and normalized into a canonical job model.