Mirantis
Senior AI Infrastructure & Platform Operations Engineer (remote in the US)
Remote, USA, United States · Senior
Sponsorship not specifiedDetected 2 days ago
AWSCloud PlatformsKubernetesLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningStakeholder ManagementNetwork EngineeringLeadershipCommunicationCollaborationMentoring
About the role
- Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites.
Responsibilities
- Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
- Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
- Lead root cause analysis activities and drive long-term corrective actions.
- Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges.
- Drive improvements in platform reliability, observability, monitoring, and operational processes.
- Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI.
- Mentor and support AI Infrastructure & Platform Operations Engineers.
- Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices.
Requirements
- 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles.
- Strong understanding of observability, monitoring, and service reliability practices.
Skills
- NVIDIA GPU infrastructure and accelerated computing platforms.
- InfiniBand networking and NVIDIA UFM.
- AI infrastructure environments.
- HPC environments.
- Platform Engineering or Site Reliability Engineering (SRE).
- Large-scale Kubernetes operations.
- Infrastructure automation technologies and Infrastructure-as-Code practices.
- Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
- Performance analysis and optimisation of distributed infrastructure platforms.
- Operate some of the most advanced AI infrastructure environments in production today.
- Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
- Help define operational standards and reliability practices for next-generation AI infrastructure services.
Apply directly at Mirantis →Create a free account for alerts like thisView Mirantis immigration profile
This listing is sourced directly from Mirantis's careers page and normalized into a canonical job model.