Lumalabs Ai

Lumalabs Ai

Software Engineer, ML Platform

Redwood City, USA

Sponsorship not specifiedDetected 337 days ago
PythonDistributed SystemsRedisAWSCloud PlatformsDockerCI/CDLinuxMachine LearningPyTorchResearch

About the role

  • We believe that multimodality is critical for intelligence.
  • You will be a foundational member of the team, designing the critical systems that allow us to train and serve next-generation AI to millions of users.
  • You will have massive ownership to: Architect end-to-end model serving pipelines and integrate new model architectures from our research team into our core, high-throughput inference engine.

Responsibilities

  • Build robust and sophisticated scheduling systems to manage jobs based on cluster availability and user priority, ensuring we optimally leverage thousands of expensive GPU resources.
  • Design and implement dynamic, traffic-based systems for hotswapping models on our GPU workers to maximize fleet efficiency and meet product SLOs.
  • Own the end-to-end CI/CD pipelines, including creating a resilient artifact store to manage all model checkpoints across multiple versions and providers.
  • Develop and maintain user-friendly APIs and interaction patterns that empower our product and research teams to ship groundbreaking features at high velocity.
  • Manage and optimize our complex inference workloads at scale, operating across multiple clusters and hardware providers.
  • About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities.
  • Where You Come In This is a rare opportunity to build the foundational infrastructure that powers our large-scale multimodal models.

Requirements

  • You are a non-negotiable fit if you have: 5+ years of professional engineering experience with deep, hands-on proficiency in Python and complex distributed systems architecture.
  • Deep expertise in our core infrastructure stack: Linux, Docker, and Kubernetes.
  • Strong experience with Redis, S3-compatible storage, and public cloud platforms (AWS).
  • Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink).
  • Familiarity with FFmpeg and multimedia processing pipelines.
  • 5+ years of professional engineering experience with deep, hands-on proficiency in Python and complex distributed systems architecture.
  • Experience with high-performance, large-scale ML systems (managing >100 GPUs).

Benefits

  • To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision.
  • What Sets You Apart (Bonus

Company info

  • So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
  • We believe that reliable, high-performance infrastructure is the single biggest differentiating factor between success and failure in achieving our mission.

This listing is sourced directly from Lumalabs Ai's careers page and normalized into a canonical job model.