Eventualcomputing
Research Engineer, Multimodal Data
San Francisco
Sponsorship not specified$60k-$100kDetected 83 days ago
SnowflakeDatabricksAWSCloud PlatformsMachine LearningSparkComputer VisionSystems EngineeringRoboticsSensors5G/LTEResearch
About the role
- Physical AI teams have raw footage, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation.
- This is a research engineering role - meaning you'll read papers and run experiments, but you ship to production and your work is judged by what it does for customer training runs.
Responsibilities
- Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale.
- Drive down per-clip annotation cost - model selection, distillation, batching, decode pipelining - so "annotate every clip in a 10K-hour corpus" stays economical.
- Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering.
- Work directly with researchers at our partner labs - your shortest feedback loop is their next training iteration.
- As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content.
- Team-building events and poker nights.
Nice to have
- ML/AI research background - papers, citations, or a research org on your resume.
- Hands-on time with big-data frameworks like Spark, Ray, or Daft.
- Worked on embeddings, retrieval, or content-aware search at scale.
- Experience designing labeling taxonomies or running annotation programs.
Compensation
- $60k-$100k
Benefits
- Competitive comp and meaningful startup equity.
- Commuter benefit.
- Health, vision, and dental coverage.
- Strong familiarity with modern vision and multimodal models - convolution nets, VLMs, VQA, embeddings - and a sense for the SOTA that's actually deployable today vs. on a leaderboard.
- Experience training vision or multimodal models from scratch (not just calling APIs).
Company info
- Every breakthrough Physical AI system - humanoid robots, autonomous vehicles, video generation models - is trained on petabytes of video, lidar, radar, and sensor data.
- But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI.
- As a result, robotics and video-AI teams iterate on model improvement about once a week.
- Most of that week isn't training - it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up.
- GPU bandwidth has grown 2-3× per generation.
- Storage and pipelines haven't.
- The gap widens every year.
- Eventual was founded in 2022 to close it.
- Our open-source engine, Daft https://daft.ai/, is the distributed data engine purpose-built for multimodal AI - already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens.
- We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop.
- Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate.
- One iteration per day becomes the norm.
- We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today.
- We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc http://Array.vc, and angels from the co-founders of Databricks and Perplexity.
Apply directly at Eventualcomputing →Create a free account for alerts like thisView Eventualcomputing immigration profile
This listing is sourced directly from Eventualcomputing's careers page and normalized into a canonical job model.