Ifm Us
Research Scientist, Agentic Data & Benchmarking
Sunnyvale, CA
Sponsorship not specified$150k-$450kDetected 43 days ago
PythonMachine LearningPyTorchData EngineeringLLMsStatisticsResearchExperimental Design
About the role
- The Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer.
- These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap.
- This is a research scientist position for someone who wants depth in data and measurement rather than breadth across the whole stack.
Responsibilities
- Design and run evaluations of agentic capabilities - multi-step reasoning, tool use, long-horizon planning, computer use, and safety - turning ambiguous notions of "intelligence" into defensible, reproducible metrics.
- Run experiments characterizing how prompting, sampling, scaffolding, and environment design affect agentic performance on internal and public benchmarks.
- Design and scale RL environments and reward signals, and measure their impact on model performance.
- Manage technical relationships with external data vendors and domain experts, evaluating data quality and iterating quickly on feedback.
- Develop QA frameworks that catch reward hacking, label noise, and contamination, keeping data and benchmark quality high.
- Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts.
- The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both.
Requirements
- 2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems).
- Strong Python and PyTorch development experience.
- Hands-on experience using LLM agents in your personal or professional work.
- 2+ years of experience with a clear emphasis on evaluations and/or training-data curation for
Nice to have
- Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts.
- Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use.
- Experience evaluating or generating data for software-engineering or computer-use agents.
- Contributions to published research, public benchmarks, and/or open-source ML software.
- Representative projects
- Diagnose a mid-training regression: an eval suite returns anomalous numbers and you determine whether it's the model, the harness, the data, or the infrastructure.
- Take a flaky distributed eval pipeline and make it reliable - better retries, better observability, faster feedback to researchers.
- We encourage you to apply even if you don't meet every qualification listed.
Skills
- trajectories, tool-use traces, and task datasets for new capabilities.
- Contribute to technical reports, research publications, and open-source benchmarks and tooling.
- Academic qualifications
- BS, MS, or PhD (or equivalent experience) in Computer Science, Machine Learning, or a related field.
Compensation
- $150k-$450k
Benefits
- Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health.
This listing is sourced directly from Ifm Us's careers page and normalized into a canonical job model.