Altruist

Altruist

Staff Back End Engineer, Evals - Hazel AI

San Francisco, CA · Staff+

Sponsorship not specified$275k-$325kDetected 56 days ago
SQLCI/CDMachine LearningdbtData EngineeringLLMsRAGTaxResearchCommunicationCollaboration

About the role

  • Architect our evaluation platform from first principles - the observability, scoring, golden datasets, verification agents, and CI/CD integration that define standards of quality.

Responsibilities

  • Design and build Hazel's evals platform end-to-end - online scoring, offline benchmarks, regression suites, LLM-as-judge pipelines, and human-in-the-loop review workflows across every Hazel surface.
  • Build production observability and monitoring for AI quality: hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals across tax planning, financial planning, investment analysis, and operational AI workflows.
  • Build and steward Hazel's golden datasets in close partnership with SMEs and a network of practicing advisors, CFPs, and tax professionals - translating their tacit expertise into precise, measurable eval criteria.
  • Develop LLM verification agents that catch hallucinations, computational errors, and compliance violations before they ever reach an advisor or client.
  • Partner with the team building Hazel's model-agnostic orchestration harness to evaluate cross-model and cross-provider performance, surface tradeoffs, and inform routing decisions across Anthropic, OpenAI, and self-hosted models.

Requirements

  • 8+ years of engineering experience, with at least 2 years focused on evaluation infrastructure, model quality, fine-tuning, or ML platform work for production systems.
  • Experience designing and curating golden datasets - sampling strategies, inter-rater agreement, dataset versioning, and managing the long tail of edge cases.
  • Comfort working across the stack - data engineering (SQL, dbt, warehouses), backend integration (APIs, async pipelines, queues), and observability tooling.
  • You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria - and explain quality tradeoffs clearly to engineers, product managers, and leadership.
  • Deep familiarity with evaluation and scoring methodologies for modern AI systems - RAG evaluation, document processing, fine-tuned model assessment, agentic and tool-use system evaluation, LLM-as-judge frameworks, and human evaluation protocols.
  • Strong communication skills. You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria - and explain quality tradeoffs clearly to engineers, product managers, and leadership.
  • A bias toward shipping. You believe great evals enable speed, not just safety, and you build tools that engineers actually want to use.
  • Bonus Points:
  • Prior experience at an applied AI company building evals, model quality, or applied research infrastructure.
  • Background in regulated industries (financial services, healthcare, legal) where accuracy, auditability, and the cost of a wrong answer are unusually high.
  • Experience building human-in-the-loop labeling workflows, annotation tooling, or red-teaming programs.
  • San Francisco, CA salary range

Nice to have

  • Experience evaluating multi-step agentic workflows, tool-use systems, or RAG pipelines in production.
  • Familiarity with frameworks like Braintrust, Langfuse or similar - including a clear point of view on when to use which.
  • Domain knowledge of wealth management, tax planning, or financial planning - or genuine excitement to learn it deeply alongside our SME bench.
  • Stunning, amenity-filled office spaces in Culver City, CA, San Francisco, CA, and Dallas, TX.
  • Financial guidance program (includes counseling on navigating debt, tracking personal spend, saving and planning goals, home-purchasing preparedness, etc.).
  • One month work from anywhere policy (with the exception of a few countries).

Skills

  • About Altruist
  • Kindness - Kindness doesn't just equal niceness.

Compensation

  • San Francisco, CA salary range

Company info

  • Define quality SLOs for each Hazel surface and build alerting that catches regressions in production before our customers do - especially for high-stakes flows like tax and financial planning.

This listing is sourced directly from Altruist's careers page and normalized into a canonical job model.