Deeter Analytics

Deeter Analytics

Head of AI Inference & MLOps

Austin Area · Senior

Sponsorship not specifiedDetected 130 days ago
KubernetesMachine LearningLLMsMLOpsProduct StrategyAccountingControls

About the role

  • We need a senior operator-builder who can sit at the intersection of:
  • That may include direct enterprise workloads, marketplace distribution, API-based reselling, model hosting, fine-tuned/private deployments, and emerging inference channels.
  • You should know what makes money on modern inference hardware, what does not, and why.

Responsibilities

  • customer / partner integration
  • Build and lead the inference monetization strategy for our first 7MW deployment and expansion to 50MW
  • Lead decisions around multi-tenant vs single-tenant deployments, reserved vs on-demand capacity, and when to prioritize direct contracts over marketplace traffic

Requirements

  • Significant experience in production AI/LLM inference, MLOps, model serving, or AI infrastructure monetization
  • Proven experience running or scaling GPU-backed inference systems in production
  • Strong understanding of modern inference runtimes, serving frameworks, and optimization techniques
  • Experience with one or more of:
  • Familiarity with AI inference aggregators, routing platforms, and model marketplaces
  • Experience designing multi-tenant GPU systems with strong isolation and predictable performance
  • Experience with advanced observability, token-level metering, cost accounting, and SLA enforcement
  • Build and manage the team required to scale this function over time

Nice to have

  • Triton Inference Server
  • Kubernetes-based GPU orchestration
  • custom routing / scheduler layers
  • Experience optimizing for real-world production metrics such as throughput, latency, GPU utilization, availability, and cost efficiency
  • Strong understanding of LLM inference economics, including tradeoffs among model size, quantization, latency, throughput, memory footprint, and customer willingness to pay
  • Ability to translate infrastructure capability into a pricing and product strategy
  • Strong technical judgment on model selection, infrastructure topology, and commercialization strategy
  • Experience monetizing large-scale NVIDIA GPU infrastructure

Skills

  • real-time inference
  • reasoning workloads
  • premium low-latency API traffic
  • batch / overflow workloads
  • dedicated enterprise deployments
  • private/fine-tuned model hosting

Compensation

  • Competitive salary, bonus, and equity participation tied to the scale, importance, and revenue generated from the role.

This listing is sourced directly from Deeter Analytics's careers page and normalized into a canonical job model.