Fabrion

Fabrion

ML Ops Engineer — Agentic AI Lab (Founding Team)

San Francisco Bay Area · Full-time

Sponsorship not specifiedDetected 345 days ago
PythonGoRustBashReactNext.jsHTMLFull-Stack DevelopmentGitVector DatabasesCloud PlatformsDockerKubernetesTerraformHelmCI/CDGitHub ActionsPrometheusGrafanaDevOpsPlatform EngineeringMachine LearningAirflowLLMs

About the role

  • Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models.
  • We're hiring an ML Ops Engineer to be the glue between ML research and production systems - responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric.
  • You'll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e

Responsibilities

  • Build and maintain secure, scalable, and automated pipelines for:
  • Containerize models and agents using Docker, with reproducible builds and CI/CD via
  • Implement and enforce model governance: versioning, metadata, lineage, reproducibility,
  • Create and manage evaluation and benchmarking frameworks (e.g. OpenLLM-Evals,
  • Support deployment of agentic apps with LangGraph, LangChain, and custom inference
  • Strong sense of ownership - prefers to build systems rather than wait for tickets

Requirements

  • Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints)

Nice to have

  • experience with SOC2, HIPAA, or GovCloud-grade model operations
  • Mistral, Falcon, Mixtral
  • Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA)
  • Implemented model-level RBAC, usage tracking, audit trails
  • Integrated with API rate limits, tenant billing, and SLA observability
  • Experience with policy-as-code systems (OPA, Rego) and access layers
  • 5+ years as a full stack or backend engineer
  • Experience owning and delivering production systems end-to-end

Skills

  • HuggingFace Hub
  • LLM fine-tuning, SFT, LoRA, RLHF, DPO training
  • RAG embedding pipelines with dynamic updates
  • Model conversion, quantization, and inference rollout
  • Integrate with security and access control layers (OPA, ABAC, Keycloak) to enforce
  • Instrument observability for model latency, token usage, performance metrics, error
  • Obsessive about reproducibility, observability, and traceability
  • Comfortable with a hybrid team of AI researchers, DevOps, and backend engineers
  • Interested in aligning ML systems to product delivery, not just papers

Compensation

  • Competitive salary + meaningful equity (founding tier)
  • Backed by 8VC, we're building a world-class team to tackle one of the industry's most critical infrastructure problems.
  • Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models.
  • We're hiring an ML Ops Engineer to be the glue between ML research and production systems - responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric.
  • You'll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e
  • Build and maintain secure, scalable, and automated pipelines for:

Benefits

  • Values equity and impact over prestige or hierarchy

This listing is sourced directly from Fabrion's careers page and normalized into a canonical job model.