Mirantis
AI Infrastructure & Platform Operations Engineer (remote in the US)
Remote, USA , United States
Sponsorship not specifiedDetected 5 days ago
Distributed SystemsAWSCloud PlatformsKubernetesLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningIncident ResponseNetwork EngineeringCommunicationCollaborationProblem Solving
About the role
- Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites.
Responsibilities
- Monitor, operate, and support production AI infrastructure platforms.
- Support NVIDIA GPU infrastructure and associated platform services.
- Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues.
- Maintain operational documentation, runbooks, and knowledge articles.
Requirements
- 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
- Working knowledge of Kubernetes in production environments.
- Experience supporting production infrastructure and services.
- Experience working within structured operational and incident management processes.
- Ability to work within a shift-based operational environment.
- Experience in one or more of the following areas is highly desirable:
Nice to have
- Preferred Experience
Skills
- NVIDIA GPU infrastructure and accelerated computing platforms.
- InfiniBand networking and NVIDIA UFM.
- Kubernetes platform operations.
- AI infrastructure or HPC environments.
- Site Reliability Engineering (SRE) or Platform Engineering.
- Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
- Infrastructure automation technologies and Infrastructure-as-Code practices.
- Large-scale distributed systems and production platforms.
- Work with some of the most advanced AI infrastructure environments in production today.
- Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
- Help define how next-generation AI infrastructure is operated and supported.
- Join a growing organisation investing heavily in AI infrastructure and platform services.
Apply directly at Mirantis →Create a free account for alerts like thisView Mirantis immigration profile
This listing is sourced directly from Mirantis's careers page and normalized into a canonical job model.