Baseten
Engineering
San Francisco
Sponsorship not specifiedDetected 13 days ago
Machine LearningSparkLLMsResearchLeadershipCollaborationPublic Speaking
About the role
- This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers.
- This role is not for someone who wants to step away from the technical work.
- Baseten Embeddings Inference: The fastest embeddings solution available https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/
Responsibilities
- Lead, grow, and mentor a team of GPU kernel engineers; own hiring, performance, and career development
- Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy
- Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams
- Establish and maintain a high technical bar for kernel quality, performance, and correctness across the team's output
- Review kernel designs and implementations with enough depth to give meaningful feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies
- Build the processes that allow a highly technical, distributed team to ship with velocity and rigor
Requirements
- You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level
- Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology
Nice to have
- Hands-on experience with Triton, CUTLASS, or CuTe DSL
- Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing
- Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem
- Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership
Skills
- Open-source contributions to GPU libraries or inference frameworks
- Experience presenting technical work at NVIDIA GTC, MLSys, or similar venues
Compensation
- Competitive compensation, including meaningful equity.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
- If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
Company info
- At Baseten, we are committed to fostering a diverse and inclusive workplace.
Apply directly at Baseten →Create a free account for alerts like thisView Baseten immigration profile
This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.