Cantina Labs

Cantina Labs

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Remote (U.S

Sponsorship not specified$200k-$220kDetected 1 day ago
C++Node.jsMachine LearningPyTorchData EngineeringNLPResearchExperimental DesignLeadership

> stay_score

odds of building a lasting career here

47Sponsors, lottery-bound
Cap-exempt (no lottery)0
Sponsors this role35
Entry-level history0
PERM / green-card track0
Lottery odds (Level IV)94
Fits your clock70

Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.

Lottery odds assume a STEM candidate.

Personalize to your clock →

> community_outcomes

No reports yet — be the first to help the next applicant.

About the role

  • In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
  • See research and engineering as two sides of the same coin and enjoy owning work end-to-end.
  • Are results-oriented, flexible, and willing to pick up whatever moves the needle.

Responsibilities

  • Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.
  • Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.
  • Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.
  • Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.
  • Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.
  • Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.
  • Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.
  • GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.
  • Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.
  • Tool Development: Develop and improve dev tooling to enhance team productivity.

Requirements

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.

Compensation

  • The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000).
  • When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
  • Competitive salary and generous company equity

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • One Medical membership, and more!

This listing is sourced directly from Cantina Labs's careers page and normalized into a canonical job model.