Deeter Analytics
Head of AI Inference & MLOps
Austin Area · Senior
Sponsorship not specifiedDetected 130 days ago
KubernetesMachine LearningLLMsMLOpsProduct StrategyAccountingControls
About the role
- We need a senior operator-builder who can sit at the intersection of:
- That may include direct enterprise workloads, marketplace distribution, API-based reselling, model hosting, fine-tuned/private deployments, and emerging inference channels.
- You should know what makes money on modern inference hardware, what does not, and why.
Responsibilities
- customer / partner integration
- Build and lead the inference monetization strategy for our first 7MW deployment and expansion to 50MW
- Lead decisions around multi-tenant vs single-tenant deployments, reserved vs on-demand capacity, and when to prioritize direct contracts over marketplace traffic
Requirements
- Significant experience in production AI/LLM inference, MLOps, model serving, or AI infrastructure monetization
- Proven experience running or scaling GPU-backed inference systems in production
- Strong understanding of modern inference runtimes, serving frameworks, and optimization techniques
- Experience with one or more of:
- Familiarity with AI inference aggregators, routing platforms, and model marketplaces
- Experience designing multi-tenant GPU systems with strong isolation and predictable performance
- Experience with advanced observability, token-level metering, cost accounting, and SLA enforcement
- Build and manage the team required to scale this function over time
Nice to have
- Triton Inference Server
- Kubernetes-based GPU orchestration
- custom routing / scheduler layers
- Experience optimizing for real-world production metrics such as throughput, latency, GPU utilization, availability, and cost efficiency
- Strong understanding of LLM inference economics, including tradeoffs among model size, quantization, latency, throughput, memory footprint, and customer willingness to pay
- Ability to translate infrastructure capability into a pricing and product strategy
- Strong technical judgment on model selection, infrastructure topology, and commercialization strategy
- Experience monetizing large-scale NVIDIA GPU infrastructure
Skills
- real-time inference
- reasoning workloads
- premium low-latency API traffic
- batch / overflow workloads
- dedicated enterprise deployments
- private/fine-tuned model hosting
Compensation
- Competitive salary, bonus, and equity participation tied to the scale, importance, and revenue generated from the role.
Apply directly at Deeter Analytics →Create a free account for alerts like thisView Deeter Analytics immigration profile
This listing is sourced directly from Deeter Analytics's careers page and normalized into a canonical job model.