Epsilon Health
Research Scientist - Post-training / RL
San Francisco, CA
Sponsorship not specifiedDetected 1 day ago
Distributed SystemsAlgorithmsMachine LearningPyTorchNLPRadiologyResearch
> stay_score
odds of building a lasting career here
16Unrated
Cap-exempt (no lottery)0
Sponsors this role0
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70
No strong sponsorship signal in the public record yet. In the full product we resolve the exact legal entity and show its filing history with a confidence score — treat as unverified until then.
Lottery odds assume a STEM candidate.
Personalize to your clock →> community_outcomes
No reports yet — be the first to help the next applicant.
About the role
- We're seeking a Research Scientist with deep expertise in post-training and reinforcement learning to join our ML Research team.
Responsibilities
- Develop inference-time methods including best-of-N sampling against reward models and grounding-aware decoding, and distill the resulting gains back into the policy.
Requirements
- Track record of implementing complex models from research papers and adapting them to new domains
- Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems
- Experience with autoregressive language modeling and instruction tuning
- Strong software engineering skills and ability to write production-quality code
- 6+ years of academia/industry experience in reinforcement learning, post-training, or multimodal machine learning
- Deep expertise in post-training large language or vision-language models (e.g., Qwen-VL, InternVL, LLaVA, or similar architectures)
- Strong foundation in modern post-training and reinforcement learning techniques including:
- Group-relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi-reward objectives
- Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals
- Reward model training: pairwise and generative reward models, outcome and process supervision
- Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF
- Inference-time compute scaling, including best-of-N sampling and verifier-guided decoding
- Practical experience diagnosing and mitigating reward hacking and reward over-optimization
- Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF
Nice to have
- Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)
- Experience with grounded generation tasks (visual grounding, referring expression comprehension)
- Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection
- Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications
Benefits
- Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity-relation matching, grounding IoU, measurement accuracy, and reporting schema compliance.
- Stay current with cutting-edge research in reinforcement learning, reward modeling, and multimodal post-training.
- Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post-training medical VLMs at scale.
Company info
- We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics.
- Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes.
- We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.
Apply directly at Epsilon Health →Create a free account for alerts like thisView Epsilon Health immigration profile
This listing is sourced directly from Epsilon Health's careers page and normalized into a canonical job model.