Montauk Capital
Head of Inference, Stealth Edge AI Co
New York City · Full-time
Sponsorship not specifiedDetected 77 days ago
C++Node.jsDistributed SystemsKubernetesLLMsSystems EngineeringLoad BalancingLeadership
About the role
- We are seeking a visionary and execution-oriented Head of Inference.
- You will be a senior, hands-on technical leader and the technical authority on inference in the room.
Responsibilities
- Create the inference strategy and define the inference architecture for Edge AI
- Own the inference serving layer end-to-end: vLLM, TensorRT-LLM, Triton, or equivalent
- Drive cost-per-token optimization
- Optimize GPU utilization, KV-cache management, and batching for production workloads
- Own observability and reliability SLAs
- Build distributed inference pipelines across multi-GPU, multi-node edge deployments
- Translate complex inference requirements for infrastructure designs
- Engage credibly with investors, partners, and technical stakeholders, represent the company externally
- Category-Defining Opportunity: Solving the AI inference bottleneck without the burden of power and infrastructure constraints Own the metro edge inference across heterogeneous, disparate compute nodes
- Studio Support: Leverage Montauk Capital's resources, network, and operational expertise during critical early stages
Requirements
- Deep knowledge and are excited about model serving, and the practical engineering required to make an inference system work on real hardware.
Nice to have
- Head of Inference
- About Montauk Capital
- About Stealth Edge AI Co
- By leveraging existing infrastructure for inference deployment, Edge AI provides low-latency, SLA-guaranteed performance across diverse GPU SKUs and colocation environments.
- Our technology intelligently routes traffic based on demand proximity and real-world network limitations, bypassing the heavy power and infrastructure requirements of traditional hyperscalers.
- Set performance baselines and SLAs for inference latency and throughput, plus observability and performance SLA's
Compensation
- True ownership over what you build
- AI spending projected to exceed hundreds of billions annually, 54GW of AI Inference demand expected by 2030
- Massive Market Opportunity: AI spending projected to exceed hundreds of billions annually, 54GW of AI Inference demand expected by 2030
Benefits
- Competitive compensation + equity: True ownership over what you build
- You can take a vision and initial concept and translate it into a viable POC quickly and are comfortable making foundational technical decisions quickly, in ambiguity, and building first of a kind.
Apply directly at Montauk Capital →Create a free account for alerts like thisView Montauk Capital immigration profile
This listing is sourced directly from Montauk Capital's careers page and normalized into a canonical job model.