Collective Health

Collective Health

Senior Software Engineer in Test (AI Agentic Systems)

Lehi, UT | Plano, TX · Senior

Sponsorship not specified$99k-$124kDetected 8 days ago
PythonSQLBigQueryCI/CDPandasLLMsRAGAgentic AILangGraphComplianceAuditingpytestHIPAACollaborationMentoring

About the role

  • You will be the quality owner for an LLM-based multi-agent pipeline that autonomously adjudicates health insurance claims for self-funded plan sponsors.
  • You will work at the intersection of Vertex AI, healthcare compliance, and high-scale data engineering.
  • Your work directly determines whether claims are paid correctly and whether the company can withstand a Department of Labor (DOL) or state DOI audit.

Responsibilities

  • Golden Set Governance: Build and maintain a versioned library of "Grounding Data" results by working with senior claims examiners to define "Ground Truth."
  • Model-as-a-Judge Automation: Design automated "LLM-grading-LLM" workflows using custom rubrics to score factual grounding and policy compliance.
  • Semantic Assertion Framework: Develop testing libraries that move beyond string matching to validate semantic equivalence and numerical accuracy in agent outputs.
  • Auto-SxS: Own the automated pairwise comparison process to detect logic drift between "New" and "Production" agent versions.
  • Mocking & Resilience: Build a Vertex AI/ADK mocking layer to simulate model responses, allowing for thousands of logic tests in seconds with zero API costs.
  • Python SDET Expertise: Expert in Python and pytest, specifically building custom mocking frameworks for external APIs ( Vertex AI/ADK ).
  • Build and maintain a versioned library of "Grounding Data" results by working with senior claims examiners to define "Ground Truth."
  • Design automated "LLM-grading-LLM" workflows using custom rubrics to score factual grounding and policy compliance.
  • Own the automated pairwise comparison process to detect logic drift between "New" and "Production" agent versions.
  • Build a Vertex AI/ADK mocking layer to simulate model responses, allowing for thousands of logic tests in seconds with zero API costs.

Requirements

  • Hands-on experience with Vertex AI Experiments, Auto-SxS, and Cloud Logging for trace analysis.
  • Ability to analyze "System Instructions" and refine prompts based on failed test cases to close logic gaps.
  • Familiarity with claims adjudication concepts (pend reason codes, COB, eligibility, stop-loss).
  • Required Skills (The Core Bar)
  • AI/LLM Observability: Hands-on experience with Vertex AI Experiments, Auto-SxS, and Cloud Logging for trace analysis.

Nice to have

  • Preferred Skills (The "Nice-to-Haves")

Skills

  • Trajectory Evaluation (The "How")
  • Use Vertex AI traces to programmatically verify that mandatory tools (via MCP) were invoked with correct arguments.
  • Expert-level SQL (BigQuery) and Pandas skills to "diff" massive datasets and identify adjudication discrepancies.

Compensation

  • The actual pay rate offered within the range will depend on factors including geographic location, qualifications, experience, and internal equity.
  • In addition to the salary, you will be eligible for 115000 stock options and benefits like health insurance, 401k, and paid time off.
  • Lehi, UT Pay Range
  • $99,200 - $124,000 USD
  • Plano, TX Pay Range
  • $109,120 - $136,400 USD

Benefits

  • At Collective Health, we're transforming how employers and their people engage with their health benefits by seamlessly integrating cutting-edge technology, compassionate service, and world-class user experience design.
  • Mission-driven culture that values innovation, collaboration, and a commitment to excellence in healthcare
  • Flexible work arrangements and a supportive work-life balance
  • Healthcare/Claims Domain: Familiarity with claims adjudication concepts (pend reason codes, COB, eligibility, stop-loss).

Equal opportunity

  • We are an equal opportunity employer and value diversity at our company.

This listing is sourced directly from Collective Health's careers page and normalized into a canonical job model.