Epsilon Health

Epsilon Health

Research Scientist - Vision-Language Modeling

San Francisco, CA

Sponsorship not specifiedDetected 264 days ago
Distributed SystemsMachine LearningPyTorchNLPRadiologyResearch

About the role

  • We're seeking a Research Scientist with deep expertise in Vision Language Modeling (VLMs) to join our ML team.
  • This role focuses on training and fine-tuning vision-language models (VLMs) that can generate accurate & grounded radiology reports across multiple imaging modalities including X-rays, CT scans, and MRI.

Responsibilities

  • Contribute hands-on to all stages of model development including dataset curation, architecture design, distributed training, post-training optimization, and production deployment.

Requirements

  • Track record of implementing complex models from research papers and adapting them to new domains
  • Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems
  • Experience with autoregressive language modeling and instruction tuning
  • Strong software engineering skills and ability to write production-quality code
  • 6+ years of academia/industry experience in vision-language modeling, multimodal learning, or related fields
  • Deep expertise in training and fine-tuning large vision-language models (e.g., LLaVA, Flamingo, CogVLM, Qwen-VL, or similar architectures)
  • Strong foundation in modern post-training techniques including:
  • Preference optimization methods (DPO, IPO, ORPO, KTO)
  • RLHF and reward modeling
  • Inference-time compute scaling and reasoning strategies
  • Constitutional AI and other alignment techniques
  • Hands-on experience with medical imaging applications, particularly radiology report generation

Nice to have

  • Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)
  • Experience with grounded generation tasks (visual grounding, referring expression comprehension)
  • Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection
  • Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications

Benefits

  • Design, train, and scale vision-language foundation models for radiology applications.
  • Pioneer grounded report generation capabilities, enabling models to spatially localize findings within medical images using bounding boxes or segmentation masks.
  • Design rigorous evaluation frameworks that assess text for medical accuracy and writing style.
  • Stay current with cutting-edge research in vision-language modeling, medical AI, and model alignment techniques.
  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for training robust medical VLMs at scale.

Company info

  • We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
  • Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
  • We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

This listing is sourced directly from Epsilon Health's careers page and normalized into a canonical job model.