Mirantis
Product Manager - AI Inference & Model Serving
Austin, TX, United States · Senior
Sponsorship not specifiedDetected 56 days ago
Distributed SystemsFull-Stack DevelopmentAWSCloud PlatformsKubernetesPlatform EngineeringMachine LearningLLMsProduct ManagementProduct StrategyCollaboration
About the role
- This role sits at the intersection of AI inference, cloud-native infrastructure, distributed systems, and performance engineering.
- You will define how NeoClouds and Enterprise customers deploy, scale, and operate production inference services while extracting maximum performance from the underlying GPU, network, and storage infrastructure.
- The scope includes serverless inference, dedicated endpoints, workload placement, autoscaling, routing, lifecycle management, observability, and full-stack performance optimization.
Responsibilities
- Own product strategy, roadmap, and lifecycle for inference and model serving, including serverless inference, dedicated endpoints, autoscaling, routing, KV cache management, and the related observability
- Lead deep technical discovery with NeoClouds, sovereign clouds, and enterprise platform teams, and translate findings into prioritized requirements and architecture direction
- Partner with engineering on system design trade-offs across runtime integration, GPU scheduling, network, storage, and serving topology, including disaggregated serving and multi-model serving
- Build the token factory foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises
- Collaborate with a world-class, distributed team committed to openness and technical excellence
Requirements
- The ideal candidate has experience with high-performance infrastructure products and understands how production systems behave under real-world load.
- 7+ years in product management, technical product management, or a senior technical role owning AI/ML and inference product(s)
- Strong understanding of production AI inference, including model serving, serverless execution, dedicated endpoints, autoscaling, routing, workload placement, observability, and reliability
Skills
- Shape the product narrative and influence go-to-market success
- Work with an established Silicon Valley leader in the cloud infrastructure industry.
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued.
- Professional development and training.
- Attend conferences and working groups.
- Customized workstation (macOS, Windows).
- A competitive compensation package with strong benefits plan and stock options.
Company info
- This person will define how customers run production model-serving workloads at scale while improving latency, throughput, utilization, reliability, cost, and operational control.
- Drive go-to-market execution: pricing and packaging, reference architectures, sizing guides, PoC playbooks, and direct engagement with customers, analysts, and ecosystem partners
Apply directly at Mirantis →Create a free account for alerts like thisView Mirantis immigration profile
This listing is sourced directly from Mirantis's careers page and normalized into a canonical job model.