Virio

Virio

*Founding Harness Engineer

San Francisco · Full-time

Sponsorship not specified$150k-$250kDetected 57 days ago
LLMsA/B TestingHubSpotResearchProblem Solving

About the role

  • The best human strategist researches the space, finds the angle, and knows your goals and constraints - but they're capped by the hours in a day and one human career's worth of expertise.
  • Crafting "perfect" content has three components: 1.
  • Research: to identify and collect the best inputs 2.

Responsibilities

  • As a Harness Engineer, you'll own the intelligence layer that sits between our AI models and product-the system prompts, tool definitions, context management, and evaluation framework that makes agents actually work in production.
  • Design the evaluation framework (LLM-as-judge evals, deterministic tests, production monitoring) that lets us ship with confidence. You'll define what success looks like and build the metrics to measure it.
  • Own prompt versioning, experimentation, and iteration. You'll A/B test prompt variations, measure their impact on agent quality, and ship improvements across our production system.
  • Collaborate with product and engineering to identify signal-where our agents are failing, where users are stuck-and translate that into prompt and architecture improvements.
  • After 90 days: You've designed and implemented our core evaluation framework.
  • After 6 months: You own the harness layer end-to-end.
  • the data feedback loop, evals, and harness design that turn a subjective judgment into something we can measure and improve - so day by day, we can get ever closer to solving content.
  • We're looking for a Harness Engineer to obsess over AI-for-writing and own that layer end to end: the data feedback loop, evals, and harness design that turn a subjective judgment into something we can measure and improve - so day by day, we can get ever closer to solving content.
  • Build and refine the system prompts, tool integrations, and context windows that shape how our models behave across our platform-you'll see how small prompt tweaks cascade into measurable product impact. (
  • Relocation support for San Francisco

Nice to have

  • background in English, literature, creative writing, or humanities)
  • Architect the abstraction layer between our product (file systems, artifacts, skills) and the model's capabilities-making complex multi-step workflows feel natural to the model and reliable to users.
  • You'll A/B test prompt variations, measure their impact on agent quality, and ship improvements across our production system.
  • After 30 days: You understand our agent stack-the system prompts, available tools, context window strategy, and current failure modes.
  • You've shipped your first prompt improvement and measured its impact.
  • Our team uses it as the source of truth for agent quality.
  • You've shipped 2-3 meaningful prompt iterations that improved measurable outputs, and you're comfortable making architecture trade-offs.
  • You've architected improvements that reduced hallucination, improved context retention, or increased model reliability.

Compensation

  • Company-wide annual off-site

Benefits

  • Meaningful equity upside as part of the founding team
  • Medical, dental, and vision insurance

This listing is sourced directly from Virio's careers page and normalized into a canonical job model.