Apptronik
Staff MLOps Engineer
Onsite - Austin, TX · Staff+
Sponsorship not specifiedDetected 22 days ago
PythonGoRustC++Code ReviewGitAWSGCPAzureCloud PlatformsDockerKubernetesCI/CDMachine LearningMLOpsLogisticsRoboticsResearchLeadershipMentoring
About the role
- This is a hands-on technical leadership role, not a management position
- Cross-Team Authority: Serve as the primary technical point of contact for Autonomy, Data Platform, and TeleOp on all matters of model lifecycle and platform contracts.
- Reproducibility: Ensure every trained model can be traced back to the exact data and code that produced it.
Responsibilities
- Apptronik is a human-centered robotics company developing AI-powered robots to support humanity in every facet of life.
- Own the technical direction for the MLOps platform - define subsystem interfaces, drive architecture decisions, and establish engineering standards for how datasets, experiments, and models move through Apptronik's systems.
- Design and operate the dataset layer end-to-end - versioning, lineage, splits, and labeling-integration handoff.
- Build and operate a first-class model registry - versioned artifacts, metadata, evaluation results, lineage, and approval workflows.
- Own the path from registered model to running inference on Apollo - packaging (ONNX, TensorRT, torch.compile), versioning on-robot, rollback, and observability of deployed policy behavior.
- Mentor mid-level and senior engineers on the MLOps team through code review, design review, and direct collaboration.
- Partner with the Training Infrastructure engineer on the cluster/platform contract, and influence research workflows across Autonomy to standardize on the platform's primitives.
- Versioning & Lineage: Design and operate the dataset layer end-to-end - versioning, lineage, splits, and labeling-integration handoff.
- Registry: Build and operate a first-class model registry - versioned artifacts, metadata, evaluation results, lineage, and approval workflows.
- On-Robot Path: Own the path from registered model to running inference on Apollo - packaging (ONNX, TensorRT, torch.compile), versioning on-robot, rollback, and observability of deployed policy behavior.
Requirements
- Deep proficiency in Python and at least one systems-level language (Go, Rust, or C++), with demonstrated ability to make and defend architectural tradeoffs in production ML platforms
- Proven experience owning and delivering an MLOps platform end-to-end - dataset lifecycle, experiment tracking, model registry, evaluation, and serving - at a company that ships models to production
- Experience defining evaluation and qualification frameworks for ML models where the cost of a regression is high (robotics, safety-critical, or production-customer-facing)
Nice to have
- Experience deploying ML models to edge or embedded targets (on-device inference, ONNX Runtime, TensorRT, robot fleets)
- Experience with RL training and evaluation infrastructure for embodied agents (rollout workers, replay buffers, sim-eval harnesses)
- Familiarity with humanoid robotics, dexterous manipulation, or teleoperation data domains
- Experience with simulation-in-the-loop evaluation (IsaacSim, MuJoCo, or equivalent)
- Familiarity with policy gating, shadow deployments, or staged rollout strategies for autonomy
- Open-source contributions to MLOps platform tooling (MLflow, BentoML, KServe, Ray Serve, etc.)
- Prolonged periods of sitting at a desk and working on a computer
- Must be able to lift 15 pounds at times
Skills
- Serving, Packaging & Deployment to Robot
Benefits
- Vision to read printed materials and a computer screen
- Our flagship humanoid robot, Apollo, is built to collaborate thoughtfully with people, starting with critical industries such as manufacturing and logistics, with future applications in healthcare, the home, and beyond.
Apply directly at Apptronik →Create a free account for alerts like thisView Apptronik immigration profile
This listing is sourced directly from Apptronik's careers page and normalized into a canonical job model.