Epsilon Health
Research Scientist - Vision-Language Modeling
San Francisco, CA
Sponsorship not specifiedDetected 264 days ago
Distributed SystemsMachine LearningPyTorchNLPRadiologyResearch
About the role
- We're seeking a Research Scientist with deep expertise in Vision Language Modeling (VLMs) to join our ML team.
- This role focuses on training and fine-tuning vision-language models (VLMs) that can generate accurate & grounded radiology reports across multiple imaging modalities including X-rays, CT scans, and MRI.
Responsibilities
- Contribute hands-on to all stages of model development including dataset curation, architecture design, distributed training, post-training optimization, and production deployment.
Requirements
- Track record of implementing complex models from research papers and adapting them to new domains
- Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems
- Experience with autoregressive language modeling and instruction tuning
- Strong software engineering skills and ability to write production-quality code
- 6+ years of academia/industry experience in vision-language modeling, multimodal learning, or related fields
- Deep expertise in training and fine-tuning large vision-language models (e.g., LLaVA, Flamingo, CogVLM, Qwen-VL, or similar architectures)
- Strong foundation in modern post-training techniques including:
- Preference optimization methods (DPO, IPO, ORPO, KTO)
- RLHF and reward modeling
- Inference-time compute scaling and reasoning strategies
- Constitutional AI and other alignment techniques
- Hands-on experience with medical imaging applications, particularly radiology report generation
Nice to have
- Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)
- Experience with grounded generation tasks (visual grounding, referring expression comprehension)
- Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection
- Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications
Benefits
- Design, train, and scale vision-language foundation models for radiology applications.
- Pioneer grounded report generation capabilities, enabling models to spatially localize findings within medical images using bounding boxes or segmentation masks.
- Design rigorous evaluation frameworks that assess text for medical accuracy and writing style.
- Stay current with cutting-edge research in vision-language modeling, medical AI, and model alignment techniques.
- Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for training robust medical VLMs at scale.
Company info
- We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
- Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
- We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.
Apply directly at Epsilon Health →Create a free account for alerts like thisView Epsilon Health immigration profile
This listing is sourced directly from Epsilon Health's careers page and normalized into a canonical job model.