Gimlet Media
Member of Technical Staff - Kernels & GPU Performance
San Francisco, CA · Staff+
Sponsorship not specifiedDetected 134 days ago
Distributed Systems
About the role
- At Gimlet, we believe every hire changes the company.
- As a an early-stage company, talent density matters more than headcount.
- The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
Responsibilities
- This role is an opportunity to help build the execution layer that transforms theoretical hardware performance into production reality.
- Build and optimize kernels that improve latency, throughput, and hardware utilization for production AI workloads
- Develop execution strategies that unlock performance across both established and emerging accelerator architectures
- Partner with compiler, runtime, and distributed systems engineers to ensure end-to-end performance optimization
Nice to have
- Improve memory efficiency, scheduling behavior, and execution characteristics across the inference stack
- Influence how heterogeneous hardware is deployed and utilized within the next generation of AI infrastructure
- Help establish performance engineering standards that shape the future of Gimlet's execution platform
- Strong software engineering fundamentals
- Experience working on performance-critical systems close to hardware
- Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs
- Experience with CUDA, Triton, CUTLASS, or other accelerator programming models
- Deep understanding of GPU execution models (warps/wavefronts, blocks, grids)
Skills
- Experience using profiling and performance analysis tools
Company info
- Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.
- The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.
- Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
- We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.
- The future of AI will require vastly more compute than exists today.
- But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough.
- The challenge is making increasingly diverse compute work together.
- Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency.
- Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.
- large-scale AI datacenters and the orchestration platform that coordinates them.
- As a an early-stage company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.
- We are focused on making increasingly diverse compute work together.
Apply directly at Gimlet Media →Create a free account for alerts like thisView Gimlet Media immigration profile
This listing is sourced directly from Gimlet Media's careers page and normalized into a canonical job model.