Preference Model
Member of Technical Staff - Research & Post-training
San Francisco · Staff+
Sponsorship not specifiedDetected 6 days ago
PythonMachine LearningPyTorchLLMsResearchCommunicationCollaborationAdaptability
About the role
- Models of the future will be able to train themselves on tasks that they are not good at.
- We are interested in investigating how far we can push the boundaries of self-directed learning.
Responsibilities
- Architect and optimize our RL training infrastructure, from training abstractions to distributed experiment management, using frameworks like Verl, OpenRLHF, or similar. Help scale our systems to handle increasingly complex research workflows.
- Design, implement, and test training environments, evaluations, and methodologies for RL agents.
- Profile and optimize training runs end-to-end, from data loading through reward computation, to maximize experiment throughput and shorten the research iteration cycle.
- Experience building and operating ML infrastructure at scale
- Have experience evaluating model outputs and building reward or evaluation signals
- Have strong systems design and communication skills
- Visa sponsorship & relocation support available
Requirements
- Experience running end-to-end LLM post-training pipelines
- Proficiency in Python and PyTorch or JAX
- Experience with at least one modern RL training framework
Compensation
- Competitive cash and equity compensation (>90th percentile)
Benefits
- Competitive cash and equity compensation (>90th percentile)
- Opportunity to work with top machine learning engineers
- Health, vision, dental, benefits
- Train and evaluate models on our proprietary RL environments to validate data quality, surface gaps in task coverage, and close the feedback loop between environment design and model capability.
Company info
- Preference Model is building automated ML research engineering.
- Existing frontier models are brittle when applied to real-world ML tasks. The present bottleneck is the lack of high-quality RL training environments. Our first step is to build RL environments that reflect real-world complexity, with diverse tasks and robust reward functions.
- Our founding team has previous experience on Anthropic's data team building data infrastructure, and datasets behind Claude. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
- Existing frontier models are brittle when applied to real-world ML tasks.
- The present bottleneck is the lack of high-quality RL training environments.
- Our first step is to build RL environments that reflect real-world complexity, with diverse tasks and robust reward functions.
- Our founding team has previous experience on Anthropic's data team building data infrastructure, and datasets behind Claude.
- We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
- Some of the best researchers have no formal ML training and gained experience building industry products.
Visa & Work Authorization
- Visa sponsorship & relocation support available
Apply directly at Preference Model →Create a free account for alerts like thisView Preference Model immigration profile
This listing is sourced directly from Preference Model's careers page and normalized into a canonical job model.