Mirantis

Mirantis

Senior AI Infrastructure & Platform Operations Engineer (remote in the US)

Remote, USA, United States · Senior

Sponsorship not specifiedDetected 2 days ago
AWSCloud PlatformsKubernetesLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningStakeholder ManagementNetwork EngineeringLeadershipCommunicationCollaborationMentoring

About the role

  • Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites.

Responsibilities

  • Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.
  • Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.
  • Lead root cause analysis activities and drive long-term corrective actions.
  • Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges.
  • Drive improvements in platform reliability, observability, monitoring, and operational processes.
  • Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI.
  • Mentor and support AI Infrastructure & Platform Operations Engineers.
  • Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices.

Requirements

  • 7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles.
  • Strong understanding of observability, monitoring, and service reliability practices.

Skills

  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • AI infrastructure environments.
  • HPC environments.
  • Platform Engineering or Site Reliability Engineering (SRE).
  • Large-scale Kubernetes operations.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Performance analysis and optimisation of distributed infrastructure platforms.
  • Operate some of the most advanced AI infrastructure environments in production today.
  • Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
  • Help define operational standards and reliability practices for next-generation AI infrastructure services.

This listing is sourced directly from Mirantis's careers page and normalized into a canonical job model.