Lumalabs Ai

Lumalabs Ai

Senior Site Reliability Engineer

Redwood City, USA · Senior

Sponsorship not specifiedDetected 229 days ago
PythonGoBashAWSCloud PlatformsKubernetesTerraformLinuxSite Reliability EngineeringMachine LearningAirflowAI OrchestrationCybersecurityComplianceResearch

About the role

  • We believe that multimodality is critical for intelligence.
  • Our SRE team is the foundation of our research and product velocity, responsible for the thousands of NVIDIA and AMD GPUs across multiple providers that power our work.

Responsibilities

  • You won't just maintain existing clusters; you will help define how our next-generation infrastructure operates.
  • Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
  • Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices in a fast-moving AI startup environment.
  • Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
  • Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure without relying on heavy operational toil.

Requirements

  • 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep Linux Mastery: You possess deep, hands-on expertise in Linux, containerized systems, and debugging low-level system performance.
  • Expert in Technologies: You have working experiencewith Terraform, Airflow, and Ray Cloud Infrastructure Expert: You have strong experience with providers like AWS or OCI.
  • Startup DNA: You are energetic and thrive in a less structured, fast-paced environment.
  • Security-Minded: You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
  • Expert in High-Performance Networking: You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.
  • Who You Are 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
  • You have strong experience with providers like AWS or OCI.
  • You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
  • You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.

Skills

  • Experience managing large-scale GPU clusters for AI/ML workloads (training or inference).
  • Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in Data Pipeline and Infrastructure
  • This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries of scale.
  • You have working experiencewith

Benefits

  • What Sets You Apart (Bonus Points) Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.

This listing is sourced directly from Lumalabs Ai's careers page and normalized into a canonical job model.