Canvas Medical

Canvas Medical

Applied AI Software Engineer

San Francisco, CA / Remote

Sponsorship not specified$300k-$400kDetected 413 days ago
PythonSQLMachine LearningData EngineeringLLMsAgentic AIEHR/EMRResearchCommunication

About the role

  • You'll be responsible for designing and running rigorous evaluation experiments that measure performance, safety, and reliability across a wide variety of clinical, operational, and financial use cases.
  • This role is ideal for someone with deep experience evaluating LLM-based agents at scale.

Responsibilities

  • Design and execute large-scale evaluation plans for LLM-based agents performing clinical documentation, scheduling, billing, communications, and general workflow automation tasks.
  • Build end-to-end test harnesses that validate model behavior under different configurations (prompt templates, context sources, tool availability, etc.).
  • Partner with clinicians to define accurate expected outcomes (gold standard) for performance comparisons in domains of clinical consequence, and partner with other subject matter experts in other non-clinical domains.
  • Deploy and maintain ongoing sampling for post-deployment governance of agent fleets.

Requirements

  • You have extensive hands-on experience evaluating LLM-based systems, including multi-agent architectures and prompt-based pipelines.
  • You are deeply familiar with foundation model APIs (OpenAI, Claude, Gemini, etc.) and how to systematically benchmark agent performance using those models in applied settings.
  • You are comfortable collaborating across engineering, product, and clinical subject matter experts.
  • Proficiency with foundation model APIs and experience orchestrating complex agent behaviors via prompts or tools.
  • Experience designing and running high-throughput evaluation pipelines, ideally including human-in-the-loop or expert-labeled benchmarks.
  • 5+ years of experience in applied machine learning or AI engineering, with a focus on evaluation and benchmarking.
  • Superlative Python engineering skills and familiarity with experiment management tools and data engineering toolsets in general including, yes, SQL and database management.
  • Familiarity with clinical or healthcare data is a strong plus.
  • Research shows that women and other minority groups might avoid applying if they don't meet 100% of the qualifications. We encourage you to apply even if you don't meet everything listed in the job posting.
  • Canvas Medical provides equal employment opportunities to all employees and applicants for employment without regard to race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Nice to have

  • Experience with reinforcement fine-tuning, model monitoring, or RLHF is a plus.

Skills

  • The server-side SDK provides extensive tools and virtually all the context necessary for excellent agent performance.

Compensation

  • $300k-$400k

Benefits

  • You are not afraid of complexity and are energized by the rigor required in healthcare deployments.

Company info

  • We're hiring an Applied AI Software Engineer to lead evaluations for agents in development and the post-deployment fleet of agents operating in Canvas to automate work for our customers.
  • Analyze results and summarize tradeoffs in clarity for product and engineering stakeholders, as well as for technical stakeholders among our customers and the broader market.
  • Your have effectively engaged your marketing counterparts to translate your work into key messages to the market and to Canvas customers.

This listing is sourced directly from Canvas Medical's careers page and normalized into a canonical job model.