Gimlet Media
Member of Technical Staff - Compilers
San Francisco, CA · Staff+
Sponsorship not specifiedDetected 134 days ago
PythonC++Machine Learning
About the role
- At Gimlet, we believe every hire changes the company.
- As a an early-stage company, talent density matters more than headcount.
- The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
Responsibilities
- This role is an opportunity to help build the execution stack that transforms modern AI workloads into efficient programs running across diverse hardware architectures.
- To learn more about the kinds of systems we build, see our work on Corsair and low-latency speculative decoding:
- Build compiler and runtime infrastructure that improves latency, throughput, and efficiency for large-scale AI inference workloads.
- Design execution strategies that intelligently partition and coordinate workloads across heterogeneous hardware.
- Develop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestration.
Requirements
- Experience implementing IR transformations, compiler passes, lowering logic, or code generation systems
- Ability to reason about execution behavior, memory systems, scheduling, and hardware efficiency
- Experience with MLIR, LLVM, XLA, TVM, Triton, or similar compiler/runtime infrastructure
- Experience optimizing ML inference or serving workloads
- Familiarity with runtime systems, kernel dispatch, launch APIs, or memory allocators
- Experience working with GPUs, AI accelerators, or heterogeneous hardware systems
Skills
- Strong software engineering skills in C++ and/or Python
Company info
- Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.
- The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
- Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
- We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
- The future of AI will require vastly more compute than exists today.
- But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough.
- The challenge is making increasingly diverse compute work together.
- Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency.
- Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
- large-scale AI datacenters and the orchestration platform that coordinates them.
- As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
- We are not building a language compiler in isolation.
- We are building the systems that determine how AI workloads are partitioned, optimized, scheduled, and executed across the next generation of AI infrastructure.
Apply directly at Gimlet Media →Create a free account for alerts like thisView Gimlet Media immigration profile
This listing is sourced directly from Gimlet Media's careers page and normalized into a canonical job model.