Lumalabs Ai
Research Engineer - Evaluations
New York, USA
Sponsorship not specifiedDetected 331 days ago
PythonDistributed SystemsCI/CDMachine LearningTensorFlowPyTorchData EngineeringLLMsManual TestingResearch
About the role
- You'll work across research, engineering, and product teams to ensure our models are measured rigorously, consistently, and in ways that directly inform development.
- Integrate evaluation signals into training loops (including reinforcement learning and reward modeling) to continuously improve model performance.
- Stay current with emerging evaluation techniques in generative AI, multimodal LLMs, and perceptual quality assessment.
Responsibilities
- We're seeking a Research Engineer to design and scale the infrastructure that powers our model evaluation efforts.
- Design and implement scalable pipelines for automated evaluation of generative models, with a focus on visual and multimodal outputs (image, video, text, audio).
- Develop novel metrics and evaluation models that capture qualities like fidelity, coherence, temporal consistency, and alignment with human intent.
- Build infrastructure for large-scale regression testing, benchmarking, and monitoring of multimodal generative models.
- Partner with model researchers to identify failure cases and build targeted evaluation harnesses.
- About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities.
- About the Role Luma is pushing the boundaries of generative AI, building tools that redefine how visual content is created.
- This role is about building the pipelines, metrics, and automated systems that close the loop between model output, evaluation, and improvement.
- Collaborate with researchers running human studies to translate human evaluation frameworks into automated or semi-automated systems.
- Maintain dashboards, reporting tools, and alerting systems to surface evaluation results to stakeholders.
Requirements
- 5+ years of experience building ML evaluation systems, model pipelines, or large-scale infrastructure.
- Proficiency in Python and ML frameworks (PyTorch, JAX, or TensorFlow).
- Familiarity with human-in-the-loop evaluation workflows and how to scale them with automation.
Nice to have
- Prior work on perceptual metrics, multimodal evaluation benchmarks, or retrieval-based evaluation.
- Background in large-scale model training or evaluation infrastructure.
- Experience designing metrics for perceptual quality Familiarity with creative media workflows (film, VFX, animation, digital art).
- Contributions to open-source evaluation libraries or benchmarks.
Skills
- Strong software engineering skills (CI/CD, testing, data pipelines, distributed systems).
Benefits
- To go beyond language models and build more aware, capable, and useful systems, the next step function change will come from vision.
- Qualifications Master's or PhD in Computer Science, Machine Learning, or a related technical field (or equivalent industry experience).
- Strong background in machine learning, with experience in generative models (diffusion, LLMs, multimodal architectures).
- Nice to Have Experience with reinforcement learning or reward modeling.
Company info
- So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
Apply directly at Lumalabs Ai →Create a free account for alerts like thisView Lumalabs Ai immigration profile
This listing is sourced directly from Lumalabs Ai's careers page and normalized into a canonical job model.