Elicit
Machine Learning Engineer
Oakland, CA (or remote within US timezones) · Principal
Sponsorship not specifiedDetected 2006 days ago
PythonMachine LearningNLPA/B TestingResearchCommunication
About the role
- Agentic harnesses for target assessment, evidence synthesis, and experiment planning that allow models to provide guarantees about their processes
Responsibilities
- building product experiences, APIs, data integrations, evaluation systems, and reliable harnesses that make language models reliably useful and trustworthy in high-stakes domains.
- optimize benchmark numbers without much connection to user workflows or product outcomes
- This is not a role for someone who only wants to develop models in isolation from user impact.
- Build a target-assessment workflow that combines literature, genetics, chemistry, clinical, regulatory, and company data into a shareable artifact.
- Build evidence-monitoring workflows that keep teams up to date through alerts, briefs, and living reports.
- Build interfaces that make it easier to inspect, trust, and correct model outputs.
- Build workflow-specific evals and quality systems that tell us whether a product change actually helped users.
- Improve extraction, reasoning, or search quality with better prompts, better system design, or finetuning when appropriate.
Requirements
- Ability to move across backend, data, and model layers as needed
- Ability to use coding assistants effectively and thoughtfully, and has adapted their workflow to become much more effective with them
- A strong software engineering background and can build end-to-end systems, not just scripts or notebooks
- Fluency with language models to reason well about prompting, retrieval, evals, failure modes, and where (and how) finetuning is or isn't worth it
- Strong product sense and likes turning fuzzy user problems into concrete things people can use
- An excitement to solve difficult, creative problems rather than narrow optimization on well-defined benchmarks
- Clear communication with product, design, domain experts, and other engineers
- To get a sense for how some of us look at applications, see this thread https://twitter.com/stuhlmueller/status/1704543826218729868. (The short version: Wherever we can, we prefer to directly evaluate work.)
- Want to build new kinds of software made possible by language models
- Are excited to use AI tools as part of your daily engineering workflow, while still applying strong judgment
Skills
- Turning messy, ambiguous research tasks into clear product experiences
- Combining language models with external tools, structured and unstructured data, and retrieval systems
- Data integrations across literature, scientific databases, customer data, and internal tools
- Evaluation systems that help us understand whether a change actually improves user outcomes
- Trust and transparency features, like source-quality signals, intermediate reasoning, and better ways to inspect and fix outputs
- Like shipping user-facing things quickly
- Enjoy working on ambiguous problems with a lot of autonomy
- Care about product quality and user trust, not just technical novelty
Compensation
- For all roles at Elicit, we use a data-backed compensation framework to make sure our salaries are market-competitive, equitable, and simple. For this role, we're targeting starting ranges of:
Benefits
- Fully covered health, dental, vision, and life insurance for you, generous coverage for the rest of your family
- Flexible vacation policy, with a minimum recommendation of 20 days/year + company holidays
- A new Mac + $1,000 budget to set up your workstation or home office in your first year, then $500 every year thereafter
- Commuter benefits, a relocation bonus, and more!
Company info
- Elicit https://elicit.com is building the reasoning layer for science and decision-making.
- We use language models to search over 125 million papers, extract data, and surface insights so that researchers, policy-makers, and industry leaders can go from questions to evidence-backed decisions in minutes.
- Today, hundreds of thousands of researchers have used Elicit to speed up literature reviews, automate systematic reviews, and explore new domains.
- As we expand our impact beyond academic research, we are laying the groundwork for ML systems that are systematic, transparent, and unbounded https://blog.elicit.com/ai-safety/ when reasoning at scale.
- To do this, Elicit is pioneering supervision of process, not outcomes https://ought.org/updates/2022-04-06-process.
- Instead of favoring large black-box models, we break complex questions down into human-legible steps and supervise the reasoning process itself.
- This approach delivers more transparent, defensible answers today and charts a safer path toward advanced AI tomorrow.
- APIs that customers can use in their own systems
- Build enterprise APIs and structured-output pipelines that plug Elicit into customers' internal systems.
- For all roles at Elicit, we use a data-backed compensation framework to make sure our salaries are market-competitive, equitable, and simple.
- In addition to working on important problems as part of a happy, productive, and positive team, we also offer great benefits (with some variation based on work location):
This listing is sourced directly from Elicit's careers page and normalized into a canonical job model.