Elly
AI Evaluation Lead
United States
Sponsorship not specified$120k-$140kDetected 4 days ago
Machine LearningData ScienceLLMsCadenceCommunication
About the role
- You will work with an AI-generated test case library and automated scoring infrastructure that is already in place.
- Your job is to make sure we are measuring the right things, interpreting what the results are telling us, and determining what needs to change to keep the system performing well as it scales.
- Evaluation complexity grows with the platform, and this role grows with it.
Responsibilities
- Our client is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice.
- You will report to our Head of Revenue & Compliance, work closely with the AI/ML team and founders, and partner with subject matter experts who provide domain judgment on complex or ambiguous cases.
- Own the criteria and calibration for when human review is triggered: defining what rises to that level, what does not, and ensuring the threshold stays well-calibrated as the platform scales.
- Partner with subject matter experts on cases that require deeper domain judgment, and incorporate their input into evaluation design.
- We aren't building another chatbot.
- We are building the Financial Answer Machine, an intelligent guide designed to help people navigate a new financial reality.
- This is a rare opportunity to join at Day Zero and architect a business designed for outsized impact and massive scale.
- A bad output here has real consequences for real people, and this role owns making sure we catch it.
Requirements
- Ability to assess whether an eval framework is measuring the right things, not just whether it is running correctly.
- Familiarity with evaluation and observability tooling.
- You have worked on AI or ML system quality in a context where outputs had real stakes.
- You are comfortable making judgment calls in ambiguous situations rather than waiting for the answer to be obvious.
- You have enough AI/ML fluency to reason about why a system is producing what it is producing, not just whether the output looks right.
- What matters is that your review is substantive rather than mechanical, and that you can have an informed conversation with those experts about what you are seeing in the data.
- You are looking for a well-defined role with stable processes.
- As an early team member, you should expect broad ownership, frequent context shifts, and a high degree of autonomy.
- We have closed an over-subscribed seed round and are looking for founding team members to help us build a bridge between the intelligence of AI and the rigid accuracy required for financial freedom.
Skills
- Assess whether current measures are detecting the right failure modes or whether new measures are needed.
- Ensure evaluation coverage keeps pace with new domain additions and model changes before they ship.
- Evolve the evaluation framework as the system grows, new domains are added, and user patterns shift.
Compensation
- Salary: 120-140k, plus early-stage option equity.
- Final compensation will depend on level, experience, location, and scope of responsibility.
Benefits
- Inflation is rising, wages are stagnant, and the traditional "retirement" model is broken.
Company info
- Underpinned by a proprietary financial system, we are turning "average" advice into personalized, multi-modal financial power.
This listing is sourced directly from Elly's careers page and normalized into a canonical job model.