Yotta Energy

Yotta Energy

GPU Cloud Platform Engineer

United States · Full-time

Sponsorship not specifiedDetected 367 days ago
GoNode.jsGitAWSGCPAzureCloud PlatformsDockerKubernetesHelmPrometheusGrafanaPlatform EngineeringRESTgRPCLLMsLoad TestingLoad BalancingResearchCommunicationCollaboration

About the role

  • We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware.
  • You will be responsible for ensuring high availability, performance, and efficiency of containerized AI workloads-ranging from LLMs to generative models-deployed in Kubernetes-based GPU clusters.
  • Please include links to any relevant projects or contributions.

Nice to have

  • ensure stable operation of compute, network, and storage systems
  • monitor and troubleshoot online issues.
  • Conduct performance testing and evaluation of multi-node GPU clusters using standard benchmarking tools to identify and resolve performance bottlenecks.
  • Deploy and orchestrate large models (e.g., LLMs, video generation models) across multi-cluster environments using Kubernetes
  • Collect key metrics such as GPU memory usage, QPS, and response latency in real time
  • configure alert mechanisms.
  • Bachelor's degree or higher in Computer Science, Software Engineering, Electronic Engineering, or related fields
  • 3+ years of experience in system engineering or DevOps.

Company info

  • Apply: careers@yottalabs.ai
  • 🧠 About Yotta Labs
  • Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world's most demanding AI workloads.
  • Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.
  • 🛠️ Role Overview
  • We are seeking a GPU Cloud Platform Engineer to join our core infrastructure team and help build the next-generation AI compute cloud.
  • In this role, you will design, deploy, and operate large-scale, multi-cluster GPU infrastructure across data centers and cloud environments.
  • If you're passionate about high-performance systems, distributed orchestration, and scaling real-world AI infrastructure, this role offers a unique opportunity to shape the backbone of our AI cloud platform.
  • 🎯 Responsibilities
  • Build and operate large-scale, high-performance GPU clusters; ensure stable operation of compute, network, and storage systems; monitor and troubleshoot online issues.
  • Deploy and orchestrate large models (e.g., LLMs, video generation models) across multi-cluster environments using Kubernetes; implement elastic scaling and cross-cluster load balancing to ensure efficient service response under high concurrency for global users.
  • Participate in the design, development, and iteration of GPU cluster scheduling and optimization systems. Define and lead Kubernetes multi-cluster configuration standards; Optimize scheduling strategies (e.g., node affinity, taints/tolerations) to improve GPU resource utilization.
  • Build a unified multi-cluster management and monitoring system to support cross-region resource monitoring, traffic scheduling, and fault failover. Collect key metrics such as GPU memory usage, QPS, and response latency in real time; configure alert mechanisms.
  • Coordinate with IDC providers for planning and deploying large-scale GPU clusters, networks, and storage infrastructure to support internal cloud platforms and external customer needs.
  • ✅ Qualifications
  • Bachelor's degree or higher in Computer Science, Software Engineering, Electronic Engineering, or related fields; 3+ years of experience in system engineering or DevOps.
  • Familiarity with the Kubernetes ecosystem; hands-on experience with tools such as kubectl, Helm, and expertise in multi-cluster deployment, upgrade, scaling, and disaster recovery.
  • Proficient in Docker and containerization technologies; knowledge of image management and cross-cluster distribution.
  • Experience with monitoring tools such as Prometheus and Grafana; Has practical experience in GPU fault monitoring and alerting.
  • Hands-on experience with cloud platforms such as AWS, GCP, or Azure; understanding of cloud-native multi-cluster architecture.
  • Experience with cluster management tools such as Ray, Slurm, KubeSphere, Rancher, Karmada is a plus.
  • Familiarity with distributed file systems such as NFS, JuiceFS, CephFS, or Lustre; ability to diagnose and resolve performance bottlenecks.
  • Understanding of high-performance communication protocols such as IB, RoCE, NVLink, and PCIe.
  • Strong communication skills, self-motivation, and team collaboration
  • 🌟 Preferred Experience
  • Experience in developing and operating MaaS platforms or large-scale model inference clusters. Proven track record of leading multi-cluster system development or performance optimization projects.

This listing is sourced directly from Yotta Energy's careers page and normalized into a canonical job model.