Elicit

Elicit

Machine Learning Engineer

Oakland, CA (or remote within US timezones) · Principal

Sponsorship not specifiedDetected 2006 days ago
PythonMachine LearningNLPA/B TestingResearchCommunication

About the role

  • Agentic harnesses for target assessment, evidence synthesis, and experiment planning that allow models to provide guarantees about their processes

Responsibilities

  • building product experiences, APIs, data integrations, evaluation systems, and reliable harnesses that make language models reliably useful and trustworthy in high-stakes domains.
  • optimize benchmark numbers without much connection to user workflows or product outcomes
  • This is not a role for someone who only wants to develop models in isolation from user impact.
  • Build a target-assessment workflow that combines literature, genetics, chemistry, clinical, regulatory, and company data into a shareable artifact.
  • Build evidence-monitoring workflows that keep teams up to date through alerts, briefs, and living reports.
  • Build interfaces that make it easier to inspect, trust, and correct model outputs.
  • Build workflow-specific evals and quality systems that tell us whether a product change actually helped users.
  • Improve extraction, reasoning, or search quality with better prompts, better system design, or finetuning when appropriate.

Requirements

  • Ability to move across backend, data, and model layers as needed
  • Ability to use coding assistants effectively and thoughtfully, and has adapted their workflow to become much more effective with them
  • A strong software engineering background and can build end-to-end systems, not just scripts or notebooks
  • Fluency with language models to reason well about prompting, retrieval, evals, failure modes, and where (and how) finetuning is or isn't worth it
  • Strong product sense and likes turning fuzzy user problems into concrete things people can use
  • An excitement to solve difficult, creative problems rather than narrow optimization on well-defined benchmarks
  • Clear communication with product, design, domain experts, and other engineers
  • To get a sense for how some of us look at applications, see this thread https://twitter.com/stuhlmueller/status/1704543826218729868. (The short version: Wherever we can, we prefer to directly evaluate work.)
  • Want to build new kinds of software made possible by language models
  • Are excited to use AI tools as part of your daily engineering workflow, while still applying strong judgment

Skills

  • Turning messy, ambiguous research tasks into clear product experiences
  • Combining language models with external tools, structured and unstructured data, and retrieval systems
  • Data integrations across literature, scientific databases, customer data, and internal tools
  • Evaluation systems that help us understand whether a change actually improves user outcomes
  • Trust and transparency features, like source-quality signals, intermediate reasoning, and better ways to inspect and fix outputs
  • Like shipping user-facing things quickly
  • Enjoy working on ambiguous problems with a lot of autonomy
  • Care about product quality and user trust, not just technical novelty

Compensation

  • For all roles at Elicit, we use a data-backed compensation framework to make sure our salaries are market-competitive, equitable, and simple. For this role, we're targeting starting ranges of:

Benefits

  • Fully covered health, dental, vision, and life insurance for you, generous coverage for the rest of your family
  • Flexible vacation policy, with a minimum recommendation of 20 days/year + company holidays
  • A new Mac + $1,000 budget to set up your workstation or home office in your first year, then $500 every year thereafter
  • Commuter benefits, a relocation bonus, and more!

Company info

  • Elicit https://elicit.com is building the reasoning layer for science and decision-making.
  • We use language models to search over 125 million papers, extract data, and surface insights so that researchers, policy-makers, and industry leaders can go from questions to evidence-backed decisions in minutes.
  • Today, hundreds of thousands of researchers have used Elicit to speed up literature reviews, automate systematic reviews, and explore new domains.
  • As we expand our impact beyond academic research, we are laying the groundwork for ML systems that are systematic, transparent, and unbounded https://blog.elicit.com/ai-safety/ when reasoning at scale.
  • To do this, Elicit is pioneering supervision of process, not outcomes https://ought.org/updates/2022-04-06-process.
  • Instead of favoring large black-box models, we break complex questions down into human-legible steps and supervise the reasoning process itself.
  • This approach delivers more transparent, defensible answers today and charts a safer path toward advanced AI tomorrow.
  • APIs that customers can use in their own systems
  • Build enterprise APIs and structured-output pipelines that plug Elicit into customers' internal systems.
  • For all roles at Elicit, we use a data-backed compensation framework to make sure our salaries are market-competitive, equitable, and simple.
  • In addition to working on important problems as part of a happy, productive, and positive team, we also offer great benefits (with some variation based on work location):

This listing is sourced directly from Elicit's careers page and normalized into a canonical job model.