Fabrion
ML Ops Engineer — Agentic AI Lab (Founding Team)
San Francisco Bay Area · Full-time
Sponsorship not specifiedDetected 345 days ago
PythonGoRustBashReactNext.jsHTMLFull-Stack DevelopmentGitVector DatabasesCloud PlatformsDockerKubernetesTerraformHelmCI/CDGitHub ActionsPrometheusGrafanaDevOpsPlatform EngineeringMachine LearningAirflowLLMs
About the role
- Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models.
- We're hiring an ML Ops Engineer to be the glue between ML research and production systems - responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric.
- You'll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e
Responsibilities
- Build and maintain secure, scalable, and automated pipelines for:
- Containerize models and agents using Docker, with reproducible builds and CI/CD via
- Implement and enforce model governance: versioning, metadata, lineage, reproducibility,
- Create and manage evaluation and benchmarking frameworks (e.g. OpenLLM-Evals,
- Support deployment of agentic apps with LangGraph, LangChain, and custom inference
- Strong sense of ownership - prefers to build systems rather than wait for tickets
Requirements
- Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints)
Nice to have
- experience with SOC2, HIPAA, or GovCloud-grade model operations
- Mistral, Falcon, Mixtral
- Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA)
- Implemented model-level RBAC, usage tracking, audit trails
- Integrated with API rate limits, tenant billing, and SLA observability
- Experience with policy-as-code systems (OPA, Rego) and access layers
- 5+ years as a full stack or backend engineer
- Experience owning and delivering production systems end-to-end
Skills
- HuggingFace Hub
- LLM fine-tuning, SFT, LoRA, RLHF, DPO training
- RAG embedding pipelines with dynamic updates
- Model conversion, quantization, and inference rollout
- Integrate with security and access control layers (OPA, ABAC, Keycloak) to enforce
- Instrument observability for model latency, token usage, performance metrics, error
- Obsessive about reproducibility, observability, and traceability
- Comfortable with a hybrid team of AI researchers, DevOps, and backend engineers
- Interested in aligning ML systems to product delivery, not just papers
Compensation
- Competitive salary + meaningful equity (founding tier)
- Backed by 8VC, we're building a world-class team to tackle one of the industry's most critical infrastructure problems.
- Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models.
- We're hiring an ML Ops Engineer to be the glue between ML research and production systems - responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric.
- You'll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e
- Build and maintain secure, scalable, and automated pipelines for:
Benefits
- Values equity and impact over prestige or hierarchy
Apply directly at Fabrion →Create a free account for alerts like thisView Fabrion immigration profile
This listing is sourced directly from Fabrion's careers page and normalized into a canonical job model.