Cantina Labs
Member of Technical Staff, Data & ML Infrastructure for Video Models
Remote (U.S · Staff+
Sponsorship not specified$200k-$260kDetected 111 days ago
PythonAWSKubernetesMachine LearningData EngineeringResearch
About the role
- This role sits at the intersection of data engineering and ML research, making it central to how we turn messy real-world data into the fuel that moves our models forward.
- Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.
- Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.
Responsibilities
- Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and DynamoDB.
- Design and run annotation workflows across platforms such as MTurk, Prolific, and Mechanical Turk, including task design, quality control, and label validation.
- Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.
- Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.
- Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.
- Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.
- Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.
- Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.
- Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.
Requirements
- Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.
Nice to have
- Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.
Compensation
- The anticipated annual base salary range for this role is between $200,000-$260,000 (€170,000-€225,000).
- When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
- Competitive salary and generous company equity
- 15 company holidays
- 2 floating holidays
- Lifestyle spending account - $500/month to use however you'd like
Benefits
- Competitive salary and generous company equity
- Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
- 42 days of paid time off, including:
- Generous parental leave & fertility support
- 401(k) retirement savings plan
- One Medical membership, and more!
Apply directly at Cantina Labs →Create a free account for alerts like thisView Cantina Labs immigration profile
This listing is sourced directly from Cantina Labs's careers page and normalized into a canonical job model.