Lumalabs Ai
Senior Site Reliability Engineer
Redwood City, USA · Senior
Sponsorship not specifiedDetected 229 days ago
PythonGoBashAWSCloud PlatformsKubernetesTerraformLinuxSite Reliability EngineeringMachine LearningAirflowAI OrchestrationCybersecurityComplianceResearch
About the role
- We believe that multimodality is critical for intelligence.
- Our SRE team is the foundation of our research and product velocity, responsible for the thousands of NVIDIA and AMD GPUs across multiple providers that power our work.
Responsibilities
- You won't just maintain existing clusters; you will help define how our next-generation infrastructure operates.
- Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
- Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices in a fast-moving AI startup environment.
- Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
- Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure without relying on heavy operational toil.
Requirements
- 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
- Deep Linux Mastery: You possess deep, hands-on expertise in Linux, containerized systems, and debugging low-level system performance.
- Expert in Technologies: You have working experiencewith Terraform, Airflow, and Ray Cloud Infrastructure Expert: You have strong experience with providers like AWS or OCI.
- Startup DNA: You are energetic and thrive in a less structured, fast-paced environment.
- Security-Minded: You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
- Expert in High-Performance Networking: You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.
- Who You Are 5+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
- You have strong experience with providers like AWS or OCI.
- You possess a working knowledge of security best practices and familiarity with compliance frameworks, such as SOC 2 and ISO.
- You have practical experience with InfiniBand, RDMA, or RoCE and understand how to optimize throughput for massive distributed training jobs.
Skills
- Experience managing large-scale GPU clusters for AI/ML workloads (training or inference).
- Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.
- Deep expertise in Data Pipeline and Infrastructure
- This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries of scale.
- You have working experiencewith
Benefits
- What Sets You Apart (Bonus Points) Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.
Apply directly at Lumalabs Ai →Create a free account for alerts like thisView Lumalabs Ai immigration profile
This listing is sourced directly from Lumalabs Ai's careers page and normalized into a canonical job model.