Yotta Energy
GPU Cloud Platform Engineer
United States · Full-time
Sponsorship not specifiedDetected 367 days ago
GoNode.jsGitAWSGCPAzureCloud PlatformsDockerKubernetesHelmPrometheusGrafanaPlatform EngineeringRESTgRPCLLMsLoad TestingLoad BalancingResearchCommunicationCollaboration
About the role
- We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware.
- You will be responsible for ensuring high availability, performance, and efficiency of containerized AI workloads-ranging from LLMs to generative models-deployed in Kubernetes-based GPU clusters.
- Please include links to any relevant projects or contributions.
Nice to have
- ensure stable operation of compute, network, and storage systems
- monitor and troubleshoot online issues.
- Conduct performance testing and evaluation of multi-node GPU clusters using standard benchmarking tools to identify and resolve performance bottlenecks.
- Deploy and orchestrate large models (e.g., LLMs, video generation models) across multi-cluster environments using Kubernetes
- Collect key metrics such as GPU memory usage, QPS, and response latency in real time
- configure alert mechanisms.
- Bachelor's degree or higher in Computer Science, Software Engineering, Electronic Engineering, or related fields
- 3+ years of experience in system engineering or DevOps.
Company info
- Apply: careers@yottalabs.ai
- 🧠 About Yotta Labs
- Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world's most demanding AI workloads.
- Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.
- 🛠️ Role Overview
- We are seeking a GPU Cloud Platform Engineer to join our core infrastructure team and help build the next-generation AI compute cloud.
- In this role, you will design, deploy, and operate large-scale, multi-cluster GPU infrastructure across data centers and cloud environments.
- If you're passionate about high-performance systems, distributed orchestration, and scaling real-world AI infrastructure, this role offers a unique opportunity to shape the backbone of our AI cloud platform.
- 🎯 Responsibilities
- Build and operate large-scale, high-performance GPU clusters; ensure stable operation of compute, network, and storage systems; monitor and troubleshoot online issues.
- Deploy and orchestrate large models (e.g., LLMs, video generation models) across multi-cluster environments using Kubernetes; implement elastic scaling and cross-cluster load balancing to ensure efficient service response under high concurrency for global users.
- Participate in the design, development, and iteration of GPU cluster scheduling and optimization systems. Define and lead Kubernetes multi-cluster configuration standards; Optimize scheduling strategies (e.g., node affinity, taints/tolerations) to improve GPU resource utilization.
- Build a unified multi-cluster management and monitoring system to support cross-region resource monitoring, traffic scheduling, and fault failover. Collect key metrics such as GPU memory usage, QPS, and response latency in real time; configure alert mechanisms.
- Coordinate with IDC providers for planning and deploying large-scale GPU clusters, networks, and storage infrastructure to support internal cloud platforms and external customer needs.
- ✅ Qualifications
- Bachelor's degree or higher in Computer Science, Software Engineering, Electronic Engineering, or related fields; 3+ years of experience in system engineering or DevOps.
- Familiarity with the Kubernetes ecosystem; hands-on experience with tools such as kubectl, Helm, and expertise in multi-cluster deployment, upgrade, scaling, and disaster recovery.
- Proficient in Docker and containerization technologies; knowledge of image management and cross-cluster distribution.
- Experience with monitoring tools such as Prometheus and Grafana; Has practical experience in GPU fault monitoring and alerting.
- Hands-on experience with cloud platforms such as AWS, GCP, or Azure; understanding of cloud-native multi-cluster architecture.
- Experience with cluster management tools such as Ray, Slurm, KubeSphere, Rancher, Karmada is a plus.
- Familiarity with distributed file systems such as NFS, JuiceFS, CephFS, or Lustre; ability to diagnose and resolve performance bottlenecks.
- Understanding of high-performance communication protocols such as IB, RoCE, NVLink, and PCIe.
- Strong communication skills, self-motivation, and team collaboration
- 🌟 Preferred Experience
- Experience in developing and operating MaaS platforms or large-scale model inference clusters. Proven track record of leading multi-cluster system development or performance optimization projects.
Apply directly at Yotta Energy →Create a free account for alerts like thisView Yotta Energy immigration profile
This listing is sourced directly from Yotta Energy's careers page and normalized into a canonical job model.