Causal
Member of Technical Staff — Data Ingestion & Quality
San Francisco · Staff+
Sponsorship not specifiedDetected 3 days ago
Machine LearningSparkData EngineeringRoboticsSensorsResearchProblem Solving
About the role
- Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.
- We look for data engineers who are excited to tackle unsolved problems.
- Data is critical to any ML model but is especially consequential for our thesis to learn physics from sensory observations.
Responsibilities
- Build petabyte-scale data pipelines (e.g. Apache Spark) that ingest each source into our storage in standardized, training-ready form, across both batch and streaming - including the orchestration, storage, and monitoring they need where shared platform infrastructure doesn't yet exist
- Design and implement automated QA checks that continuously measure and monitor data quality over time, and own the verdicts they produce
- Write technical requirements and provide actionable feedback to external data vendors and partners
- Collaborate with researchers to validate that new and improved datasets translate into model performance
- Demonstrated experience building large-scale data pipelines, QA systems, or evaluation workflows (e.g. Spark, Ray, Beam)
- Experience working with external data vendors and partners, from technical evaluation to ongoing feedback
- Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution
Requirements
- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
- Detail-oriented in identifying subtle data inconsistencies and issues that could affect quality, with the ability to understand how quality impacts model performance
- We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.
Benefits
- Develop quality metrics that measure coverage, correctness, and consistency across sources - and catch the subtle inconsistencies (sensor bias, drift, processing artifacts) that silently degrade models
Company info
- Your mission is to own every dataset end to end - from discovering the source and securing access, to writing the pipelines that ingest it, to guaranteeing it enters training clean, standardized, and correct.
- What we're looking for
This listing is sourced directly from Causal's careers page and normalized into a canonical job model.