TrueFoundry
Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)
San Mateo, San Francisco Bay Area · Senior
Sponsorship not specifiedDetected 2 days ago
Distributed SystemsVector DatabasesKubernetesCI/CDMachine LearningLLMsRAGAgentic AILLMOpsLangGraphAI OrchestrationAccountingCustomer SupportLeadership
About the role
- The Problem We're Solving Companies are moving beyond simple chatbots to production agentic systems.
Responsibilities
- Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.
- Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform's LLMops flows.
- Build and improve tracing, benchmarking and observability for LLMs and agents - token/cost accounting, latency p95, throughput, and correctness checks.
- Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.
- Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).
- 5-10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms.
- Proven track record building observability, cost controls, and policy enforcement for production services.
- Experience building or contributing to open-source LLM orchestration tools (LangGraph, LangChain, or similar).
- This includes building robust orchestration for multi-step agents (graph/stateful workflows), model/routing logic, observability and policy enforcement (cost, data residency, rate limiting), and integrating upstream tooling like LangGraph, LangChain, vector stores, and specialized LLM runtimes.
Requirements
- Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines).
- Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design.
Nice to have
- Preferred / differentiators
Skills
- About TrueFoundry
- A way to orchestrate agents and enforce governance.
- A unified compute layer to run it all.
- That infrastructure layer is being built right now.
- The Problem We're Solving
- Companies are moving beyond simple chatbots to production agentic systems.
- These systems route between OpenAI, Anthropic, Google, and self-hosted models.
- They integrate dozens of tools via protocols like MCP.
- They orchestrate multi-agent workflows where agents coordinate with other agents.
- The infrastructure to support this doesn't exist yet.
- Intelligent routing with observability, cost policies, and fallback logic
- Centralized tool and MCP server management with security and lifecycle controls
Benefits
- Join a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure.
- Comprehensive health insurance for you and your family, including medical, dental, and vision coverage.
- 401(k) retirement plan.
- Flexible hybrid work, 2 days a week in the office (Tuesday & Wednesday), with flexibility around the schedule.
Company info
- Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.
- You'll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on TrueFoundry.
- This is a customer-facing role where you'll work closely with TrueFoundry's enterprise customers.
Apply directly at TrueFoundry →Create a free account for alerts like thisView TrueFoundry immigration profile
This listing is sourced directly from TrueFoundry's careers page and normalized into a canonical job model.