Preference Model
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco · Staff+
Sponsorship not specifiedDetected 5 days ago
Distributed SystemsAWSGCPCloud PlatformsKubernetesMachine LearningPyTorchData EngineeringNLPLLMsA/B TestingTest AutomationResearch
About the role
- Frontier research moves only as fast as its infrastructure permits.
Responsibilities
- Design, build, and scale the compute, scheduling, and data infrastructure that powers post-training research on our in-house RL environments
- Develop and maintain core ML framework primitives and internal tooling that researchers rely on daily, accelerating reproducible experimentation and reducing time from idea to result
- Build evaluation and benchmarking infrastructure, monitoring, logging, and debugging tooling, and automated testing and deployment systems, so failures are caught early and infrastructure stays reliable as it scales
- Partner directly with Research Engineers to translate research needs into infrastructure requirements, and ship fast in response to their feedback
- Visa sponsorship & relocation support available
Requirements
- Have strong software engineering fundamentals, experience building production-grade infrastructure (ideally for ML or data-intensive systems), and proficiency in core ML frameworks such as PyTorch or JAX
- Have experience with data engineering tools and building robust, scalable data pipelines
Nice to have
- Opportunity to work alongside senior and staff engineers from frontier labs and infrastructure companies, plus top ML engineers
Compensation
- Competitive cash and equity compensation (>90th percentile)
Benefits
- Competitive cash and equity compensation (>90th percentile)
- Health, vision, dental, benefits
Company info
- Preference Model is automating ML engineering and a critical component is models' abilities to develop software.
- The way we build software is changing fast. Five years ago we wrote every line of code by hand. Today, we don't. What does our work look like five years from now? We are shaping this future.
- Recent models work well on narrow tasks but are still brittle on real software work: large codebases with real conventions and technical debt, judgment-heavy design decisions, and multi-step problems. The bottleneck on fixing that is the supply of hard, high-fidelity scenarios that find where the best models still break. That is what we build.
- Our founding team has previous experience on Anthropic's data team building data infrastructure, and datasets behind Claude. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
- The way we build software is changing fast.
- Five years ago we wrote every line of code by hand.
- Today, we don't.
- What does our work look like five years from now?
- We are shaping this future.
- Recent models work well on narrow tasks but are still brittle on real software work: large codebases with real conventions and technical debt, judgment-heavy design decisions, and multi-step problems.
- The bottleneck on fixing that is the supply of hard, high-fidelity scenarios that find where the best models still break.
- That is what we build.
- Our founding team has previous experience on Anthropic's data team building data infrastructure, and datasets behind Claude.
- We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
Visa & Work Authorization
- Visa sponsorship & relocation support available
Apply directly at Preference Model →Create a free account for alerts like thisView Preference Model immigration profile
This listing is sourced directly from Preference Model's careers page and normalized into a canonical job model.