TrueFoundry

TrueFoundry

Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)

San Mateo, San Francisco Bay Area · Senior

Sponsorship not specifiedDetected 2 days ago
Distributed SystemsVector DatabasesKubernetesCI/CDMachine LearningLLMsRAGAgentic AILLMOpsLangGraphAI OrchestrationAccountingCustomer SupportLeadership

About the role

  • The Problem We're Solving Companies are moving beyond simple chatbots to production agentic systems.

Responsibilities

  • Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.
  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform's LLMops flows.
  • Build and improve tracing, benchmarking and observability for LLMs and agents - token/cost accounting, latency p95, throughput, and correctness checks.
  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.
  • Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).
  • 5-10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms.
  • Proven track record building observability, cost controls, and policy enforcement for production services.
  • Experience building or contributing to open-source LLM orchestration tools (LangGraph, LangChain, or similar).
  • This includes building robust orchestration for multi-step agents (graph/stateful workflows), model/routing logic, observability and policy enforcement (cost, data residency, rate limiting), and integrating upstream tooling like LangGraph, LangChain, vector stores, and specialized LLM runtimes.

Requirements

  • Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines).
  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design.

Nice to have

  • Preferred / differentiators

Skills

  • About TrueFoundry
  • A way to orchestrate agents and enforce governance.
  • A unified compute layer to run it all.
  • That infrastructure layer is being built right now.
  • The Problem We're Solving
  • Companies are moving beyond simple chatbots to production agentic systems.
  • These systems route between OpenAI, Anthropic, Google, and self-hosted models.
  • They integrate dozens of tools via protocols like MCP.
  • They orchestrate multi-agent workflows where agents coordinate with other agents.
  • The infrastructure to support this doesn't exist yet.
  • Intelligent routing with observability, cost policies, and fallback logic
  • Centralized tool and MCP server management with security and lifecycle controls

Benefits

  • Join a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure.
  • Comprehensive health insurance for you and your family, including medical, dental, and vision coverage.
  • 401(k) retirement plan.
  • Flexible hybrid work, 2 days a week in the office (Tuesday & Wednesday), with flexibility around the schedule.

Company info

  • Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.
  • You'll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on TrueFoundry.
  • This is a customer-facing role where you'll work closely with TrueFoundry's enterprise customers.

This listing is sourced directly from TrueFoundry's careers page and normalized into a canonical job model.