Gimlet Media

Gimlet Media

Member of Technical Staff - ML Systems & Inference

San Francisco, CA · Staff+

Sponsorship not specifiedDetected 134 days ago
PythonC++Distributed SystemsMachine LearningLLMs

About the role

  • At Gimlet, we believe every hire changes the company.
  • As a an early-stage company, talent density matters more than headcount.
  • The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.

Responsibilities

  • This role is an opportunity to help build the systems that determine how modern AI workloads are executed in production.
  • Build and optimize inference systems that improve latency, throughput, and efficiency for production AI workloads
  • Design execution strategies that intelligently balance batching, scheduling, concurrency, and resource utilization
  • Partner with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance improvements

Requirements

  • Comfort reasoning about performance, memory usage, and system behavior under load
  • Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems
  • Experience with batching, scheduling, and concurrency control in inference systems
  • Familiarity with KV cache management and memory placement strategies
  • Experience profiling and tuning latency- and throughput-critical systems

Company info

  • Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.
  • The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
  • Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
  • We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
  • The future of AI will require vastly more compute than exists today.
  • But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough.
  • The challenge is making increasingly diverse compute work together.
  • Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency.
  • Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
  • large-scale AI datacenters and the orchestration platform that coordinates them.
  • As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
  • We are not training models or tuning benchmarks in isolation.
  • We are building the systems that determine how AI workloads are served, optimized, and executed across the next generation of AI infrastructure.

This listing is sourced directly from Gimlet Media's careers page and normalized into a canonical job model.