Cantina Labs

Cantina Labs

Member of Technical Staff, Data & ML Infrastructure for Video Models

Remote (U.S · Staff+

Sponsorship not specified$200k-$260kDetected 111 days ago
PythonAWSKubernetesMachine LearningData EngineeringResearch

About the role

  • This role sits at the intersection of data engineering and ML research, making it central to how we turn messy real-world data into the fuel that moves our models forward.
  • Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.
  • Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.

Responsibilities

  • Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and DynamoDB.
  • Design and run annotation workflows across platforms such as MTurk, Prolific, and Mechanical Turk, including task design, quality control, and label validation.
  • Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.
  • Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.
  • Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.
  • Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.
  • Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.
  • Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.
  • Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.

Requirements

  • Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.

Nice to have

  • Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.

Compensation

  • The anticipated annual base salary range for this role is between $200,000-$260,000 (€170,000-€225,000).
  • When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
  • Competitive salary and generous company equity
  • 15 company holidays
  • 2 floating holidays
  • Lifestyle spending account - $500/month to use however you'd like

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • One Medical membership, and more!

This listing is sourced directly from Cantina Labs's careers page and normalized into a canonical job model.