Mirantis

Mirantis

AI Infrastructure & Platform Operations Engineer (remote in the US)

Remote, USA , United States

Sponsorship not specifiedDetected 5 days ago
Distributed SystemsAWSCloud PlatformsKubernetesLinuxPrometheusGrafanaSite Reliability EngineeringPlatform EngineeringMachine LearningIncident ResponseNetwork EngineeringCommunicationCollaborationProblem Solving

About the role

  • Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites.

Responsibilities

  • Monitor, operate, and support production AI infrastructure platforms.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Collaborate with engineering teams, hardware vendors, Data Center personnel, and service delivery teams to resolve technical issues.
  • Maintain operational documentation, runbooks, and knowledge articles.

Requirements

  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Experience working within structured operational and incident management processes.
  • Ability to work within a shift-based operational environment.
  • Experience in one or more of the following areas is highly desirable:

Nice to have

  • Preferred Experience

Skills

  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Kubernetes platform operations.
  • AI infrastructure or HPC environments.
  • Site Reliability Engineering (SRE) or Platform Engineering.
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Large-scale distributed systems and production platforms.
  • Work with some of the most advanced AI infrastructure environments in production today.
  • Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
  • Help define how next-generation AI infrastructure is operated and supported.
  • Join a growing organisation investing heavily in AI infrastructure and platform services.

This listing is sourced directly from Mirantis's careers page and normalized into a canonical job model.