Montauk Capital

Montauk Capital

Head of Inference, Stealth Edge AI Co

New York City · Full-time

Sponsorship not specifiedDetected 77 days ago
C++Node.jsDistributed SystemsKubernetesLLMsSystems EngineeringLoad BalancingLeadership

About the role

  • We are seeking a visionary and execution-oriented Head of Inference.
  • You will be a senior, hands-on technical leader and the technical authority on inference in the room.

Responsibilities

  • Create the inference strategy and define the inference architecture for Edge AI
  • Own the inference serving layer end-to-end: vLLM, TensorRT-LLM, Triton, or equivalent
  • Drive cost-per-token optimization
  • Optimize GPU utilization, KV-cache management, and batching for production workloads
  • Own observability and reliability SLAs
  • Build distributed inference pipelines across multi-GPU, multi-node edge deployments
  • Translate complex inference requirements for infrastructure designs
  • Engage credibly with investors, partners, and technical stakeholders, represent the company externally
  • Category-Defining Opportunity: Solving the AI inference bottleneck without the burden of power and infrastructure constraints Own the metro edge inference across heterogeneous, disparate compute nodes
  • Studio Support: Leverage Montauk Capital's resources, network, and operational expertise during critical early stages

Requirements

  • Deep knowledge and are excited about model serving, and the practical engineering required to make an inference system work on real hardware.

Nice to have

  • Head of Inference
  • About Montauk Capital
  • About Stealth Edge AI Co
  • By leveraging existing infrastructure for inference deployment, Edge AI provides low-latency, SLA-guaranteed performance across diverse GPU SKUs and colocation environments.
  • Our technology intelligently routes traffic based on demand proximity and real-world network limitations, bypassing the heavy power and infrastructure requirements of traditional hyperscalers.
  • Set performance baselines and SLAs for inference latency and throughput, plus observability and performance SLA's

Compensation

  • True ownership over what you build
  • AI spending projected to exceed hundreds of billions annually, 54GW of AI Inference demand expected by 2030
  • Massive Market Opportunity: AI spending projected to exceed hundreds of billions annually, 54GW of AI Inference demand expected by 2030

Benefits

  • Competitive compensation + equity: True ownership over what you build
  • You can take a vision and initial concept and translate it into a viable POC quickly and are comfortable making foundational technical decisions quickly, in ambiguity, and building first of a kind.

This listing is sourced directly from Montauk Capital's careers page and normalized into a canonical job model.