Clera
Staff Engineer — Agentic AI
San Francisco · Staff+
No sponsorship$160k-$250kDetected 19 hours ago
PythonCode ReviewLLMsAgentic AIUX ResearchMechanical DesignResearch
About the role
- Our product sits at the intersection of applied agentic AI, enterprise software automation, and deep domain expertise in CAD, simulation, and PLM.
- This role is on-site in San Francisco, CA.
Responsibilities
- Drive agent task success rate. Own the metric that matters most - can the agent actually complete the workflows engineers need? Define the eval framework, establish baselines, and systematically improve performance.
- Build rigorous evaluation infrastructure. Design benchmarks grounded in real user stories, not synthetic tasks. Think SWE-bench-level rigor applied to engineering workflows - reproducible, adversarial, and tied to actual customer value.
- Lead user story mapping and validation. Work directly with the user researcher and domain expert contractors to interview engineers, document their workflows in detail, and validate that what you're building against reflects reality - not assumptions.
- Own the agent architecture. Make foundational decisions around tool-calling strategies, state management across multi-step workflows, error recovery, model routing, and context management.
- Lead a team as a player-coach. Set technical direction, review architecture decisions, unblock the team, raise the engineering bar - and still write production code.
- 7+ years in software engineering, including at least 2 years building agentic LLM-based systems that call tools, manage multi-step workflows, handle failures, and operate under cost constraints.
- Drive agent task success rate.
Requirements
- Deep experience with LLM application architecture: model selection, context window management, retrieval strategies, tool-calling frameworks, and orchestration.
- familiarity with benchmarks such as SWE-bench, GAIA, or τ-bench.
- Proven track record of shipping AI systems with measurable outcomes (e.g., agent task success rate, cost efficiency) - not just demos.
- Strong Python skills and working knowledge of the LLM tooling ecosystem: function calling, tool use APIs, tracing/observability tools (e.g., Logfire, LangSmith), and evaluation frameworks.
- Experience leading a small technical team (3-6 engineers): setting technical direction, performing code reviews, and driving architecture decisions.
Nice to have
- Published work or open-source contributions in agentic AI systems.
- Experience with desktop automation, COM, or programmatic control of applications (beyond web APIs).
- Background in mechanical engineering, CAD/CAE, PLM, or adjacent industries.
- Familiarity with enterprise deployment constraints - running agents on locked-down corporate workstations.
- Direct line to the CTO and outsized impact at a critical stage of company growth
- Small, high-caliber team working on a technically deep problem in a largely untapped industry vertical
- Candidates must be authorized to work in the United States - visa sponsorship is not available.
Compensation
- Salary: $160,000 - $250,000 per year, depending on experience
Benefits
- Expand workflow coverage.
- Early-stage equity in a Series A company with strong investor backing
Company info
- Collaborate cross-functionally with integrations, product, and enterprise customers during POCs to align agent behavior with real-world usage.
Visa & Work Authorization
- Visa sponsorship is not available.
This listing is sourced directly from Clera's careers page and normalized into a canonical job model.