Runpod

Runpod

Manager, HPC Storage Engineer

Remote - USA

Sponsorship not specified$150k-$240kDetected 175 days ago
Cloud PlatformsLinuxSite Reliability EngineeringMachine LearningIncident ResponseLeadershipCommunicationCollaboration

About the role

  • Runpod is pioneering the future of AI and machine learning, offering cutting-edge cloud infrastructure for full‑stack AI applications.
  • As AI workloads continue to push the limits of throughput, latency, and parallelism, Runpod is investing heavily in next-generation storage architectures purpose-built for GPU-centric compute.
  • This role requires deep technical fluency and architectural leadership, combined with strong people management and operational discipline.

Responsibilities

  • Own Distributed Storage Architecture: Define, evolve, and operate Runpod's global storage platforms, supporting training, inference, checkpointing, and dataset access at scale.
  • Build the Storage Engineering Team: Manage and grow a team of storage and systems engineers. Set clear ownership, technical direction, and operational standards across regions.
  • High-Performance Shared Filesystems: Design and operate large-scale SAN and NFS deployments, including performance-sensitive shared storage for GPU clusters.=
  • End-to-End Performance Ownership: Drive performance optimization from NAND and NVMe media through controllers, networking, and client access patterns.
  • Automation & Observability: Build automation for provisioning, expansion, upgrades, and monitoring. Ensure deep observability into throughput, latency, and error characteristics.
  • Cross-Functional Collaboration: Partner with Datacenter Networking, GPU Platform, SRE, and Product teams to ensure storage systems meet evolving workload and customer needs.
  • Vendor & Partner Management: Own technical relationships with storage vendors, hardware partners, and colocation providers
  • drive roadmap alignment and issue resolution.
  • Vendor & Partner Management: Own technical relationships with storage vendors, hardware partners, and colocation providers; drive roadmap alignment and issue resolution.
  • Manage and grow a team of storage and systems engineers.

Requirements

  • Distributed Storage Expertise: 8+ years designing and operating large-scale storage systems, including SAN and NFS architectures at multi-petabyte scale.
  • Linux Systems Expertise: Strong Linux internals knowledge, including filesystems, I/O scheduling, memory management, and tuning for performance workloads.
  • Lead deployments and operations of VAST Data and experience with Lustre or similar parallel filesystems used in HPC and AI environments.
  • Hands-on experience deploying, operating, or deeply integrating VAST Data in production environments is required.
  • Experience with Lustre or comparable HPC filesystems (e.g., GPFS, BeeGFS) supporting high-concurrency workloads.
  • Advanced Filesystems & Platforms: Lead deployments and operations of VAST Data and experience with Lustre or similar parallel filesystems used in HPC and AI environments.

Nice to have

  • Experience supporting AI training pipelines, large-scale model checkpointing, and dataset streaming workloads.
  • Familiarity with RDMA fabrics and close collaboration with datacenter networking teams.
  • Experience designing storage systems for multi-tenant isolation and secure data access.
  • Background in hyperscale, HPC, or AI-focused infrastructure environments.
  • What You'll Receive:
  • Proven experience with NFS over RDMA, RDMA-capable transports, or similar technologies.
  • Familiarity with GPU Direct Storage strongly preferred.
  • High-Performance Data Paths: Proven experience with NFS over RDMA, RDMA-capable transports, or similar technologies.

Skills

  • Engineering Leadership Experience: 3+ years managing storage, systems, or infrastructure engineering teams in production environments.
  • VAST Data Experience: Hands-on experience deploying, operating, or deeply integrating VAST Data in production environments is required.
  • Parallel Filesystems: Experience with Lustre or comparable HPC filesystems (e.g., GPFS, BeeGFS) supporting high-concurrency workloads.
  • Low-Level Storage Knowledge: Deep understanding of NAND, NVMe, PCIe, storage controllers, and performance characteristics across the stack.

Compensation

  • The competitive base pay for this position ranges from $150,000 - $240,000 USD.

Benefits

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.

This listing is sourced directly from Runpod's careers page and normalized into a canonical job model.