OpenAI
Hardware Technical Program Manager, Infrastructure Partner Operations
San Francisco · Contract
Sponsorship not specifiedDetected 2 days ago
AWSGCPAzureCloud PlatformsIncident ResponseAccount ManagementSupply ChainProcess ImprovementCustomer SuccessHardware DesignResearchCommunicationPublic Speaking
About the role
- As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical.
Responsibilities
- Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations.
- Develop operational governance frameworks with strategic partners, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
- Build dashboards and reporting mechanisms that provide clear visibility into partner operational performance, risks, trends, and areas requiring executive attention.
- Drive cross-functional coordination between OpenAI teams and external infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
- Lead operational escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication.
- Establish repeatable operating rhythms with external partners, including weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives.
- Partner with Capacity Planning, Hardware Operations, Networking, Deployment, Reliability Engineering, and Supply Chain teams to ensure external infrastructure providers remain aligned with OpenAI's operational priorities.
- Identify systemic operational risks across partner organizations and proactively drive corrective actions that improve long-term operational effectiveness.
Requirements
- 7+ years of experience in Technical Program Management, Infrastructure Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments.
- Strong understanding of hyperscale cloud infrastructure, data center operations, infrastructure delivery, or large-scale distributed systems.
- Bachelor's degree in Engineering, Computer Science, Information Systems, Operations, or equivalent practical experience.
- Strong understanding of service-level agreements (SLAs), operational KPIs, incident management, escalation processes, root cause analysis, and continuous service improvement methodologies.
- Familiarity with infrastructure operations supporting GPU infrastructure, AI infrastructure, high-performance computing (HPC), or hyperscale data center environments.
- Experience managing operational relationships with external infrastructure providers, cloud service providers, hardware vendors, or strategic technology partners.
- Experience developing operational KPIs, SLAs, service health metrics, dashboards, and executive reporting for complex technical organizations.
- Demonstrated success leading cross-functional operational programs involving both internal stakeholders and external partners.
- Strong program management skills with the ability to drive accountability across organizations without direct authority.
- Excellent written and verbal communication skills with experience presenting operational performance to senior technical and executive leadership.
- Experience managing cloud infrastructure operations within organizations such as Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), or other hyperscale cloud providers.
- Experience leading operational governance, service delivery, customer success engineering, technical account management, or infrastructure operations for enterprise cloud customers.
- Experience building executive dashboards, operational scorecards, business review frameworks, and data-driven performance reporting.
Nice to have
- Preferred Skills
Benefits
- Define, track, and continuously improve key operational metrics related to infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.
Company info
- The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI's largest-scale AI systems.
- We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads.
- Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI's rapidly expanding compute environment.
- We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241.
This listing is sourced directly from OpenAI's careers page and normalized into a canonical job model.