Epsilon Health

Epsilon Health

Research Scientist - VLM Pretraining

San Francisco, CA

Sponsorship not specifiedDetected 1 day ago
Distributed SystemsMachine LearningPyTorchNLPRadiologyResearch

> stay_score

odds of building a lasting career here

16Unrated
Cap-exempt (no lottery)0
Sponsors this role0
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70

No strong sponsorship signal in the public record yet. In the full product we resolve the exact legal entity and show its filing history with a confidence score — treat as unverified until then.

Lottery odds assume a STEM candidate.

Personalize to your clock →

> community_outcomes

No reports yet — be the first to help the next applicant.

About the role

  • We're seeking a Research Scientist with deep expertise in large-scale vision-language pretraining to join our ML Research team.

Responsibilities

  • Build and tune multimodal pretraining mixtures across captioning, VQA, grounding, and retrieval tasks, balancing data sources to avoid regressions in language capability.
  • Own pretraining evaluation (zero- and few-shot transfer, probing, and downstream fine-tunability) - as the signal for base model quality.
  • Contribute hands-on to all stages of pretraining including dataset curation, architecture design, distributed training, and handoff of base checkpoints to post-training.
  • Post-training and RL are owned by a partner role you'll collaborate with closely.

Requirements

  • Experience with fine-grained visual grounding (referring expression comprehension, phrase grounding, box or mask prediction)
  • Track record of implementing complex models from research papers and adapting them to new domains
  • Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems
  • Experience with autoregressive language modeling and long-context training
  • Strong software engineering skills and ability to write production-quality code
  • 6+ years of academia/industry experience in vision-language modeling, multimodal learning, or related fields
  • Deep expertise in pretraining large vision-language models (e.g., LLaVA, Flamingo, CogVLM, Qwen-VL, InternVL, or similar architectures)
  • Strong foundation in modern VLM pretraining techniques including:
  • Vision-language connector and fusion architectures (projection, cross-attention, resampler-based)
  • Variable and high-resolution image handling (native resolution, dynamic tiling, token compression)
  • Contrastive and generative objectives for learning joint vision-language embedding spaces
  • Data and task mixture design, including curriculum and mixture-ratio ablations
  • Hands-on experience with medical imaging applications, particularly radiology report generation

Nice to have

  • Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)
  • Experience with interleaved image-text pretraining and synthetic recaptioning pipelines
  • Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection
  • Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications

Benefits

  • Design, train, and scale vision-language foundation models for radiology applications, owning the pretraining stage end to end.
  • Develop VLM architectures suited to medical imaging, including native and variable resolution handling, high-resolution tiling, connector design, and token budgets for volumetric studies.
  • Develop fine-grained visual grounding during pretraining, enabling models to localize findings within medical images using bounding boxes or segmentation masks.
  • Train joint vision-language embedding spaces using contrastive and generative objectives, including region- and sentence-level alignment between images and reports.
  • Stay current with cutting-edge research in vision-language modeling and large-scale multimodal pretraining.
  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for pretraining medical VLMs at scale.

Company info

  • We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
  • Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
  • We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

This listing is sourced directly from Epsilon Health's careers page and normalized into a canonical job model.