Baseten

Baseten

Engineering

San Francisco

Sponsorship not specifiedDetected 13 days ago
Machine LearningSparkLLMsResearchLeadershipCollaborationPublic Speaking

About the role

  • This is a player-coach role for someone who has spent years hands-on writing kernels and is now ready to multiply their impact by leading a team of elite GPU engineers.
  • This role is not for someone who wants to step away from the technical work.
  • Baseten Embeddings Inference: The fastest embeddings solution available https://www.baseten.co/blog/introducing-baseten-embeddings-inference-bei/

Responsibilities

  • Lead, grow, and mentor a team of GPU kernel engineers; own hiring, performance, and career development
  • Partner closely with the Chief Scientist, VP Engineering, and peer engineering leads to align kernel work with Baseten's broader inference stack strategy
  • Drive cross-functional collaboration between the kernel team and Model Performance, Capacity, and Infrastructure teams
  • Establish and maintain a high technical bar for kernel quality, performance, and correctness across the team's output
  • Review kernel designs and implementations with enough depth to give meaningful feedback on GPU architecture decisions, memory hierarchy tradeoffs, and optimization strategies
  • Build the processes that allow a highly technical, distributed team to ship with velocity and rigor

Requirements

  • You have written and shipped production CUDA kernels and can credibly engage with your team's work at a technical level
  • Strong understanding of GPU architecture fundamentals: memory hierarchy, warp execution, tensor cores, occupancy tradeoffs, and profiling methodology

Nice to have

  • Hands-on experience with Triton, CUTLASS, or CuTe DSL
  • Background in LLM inference kernels: attention variants, GEMMs, quantization (FP8/FP4), MoE routing
  • Experience with NVIDIA GPU architectures (Hopper or Blackwell preferred) and the CUDA ecosystem
  • Demonstrated ability to set technical direction, prioritize a roadmap, and communicate clearly across engineering and leadership

Skills

  • Open-source contributions to GPU libraries or inference frameworks
  • Experience presenting technical work at NVIDIA GTC, MLSys, or similar venues

Compensation

  • Competitive compensation, including meaningful equity.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
  • If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

Company info

  • At Baseten, we are committed to fostering a diverse and inclusive workplace.

This listing is sourced directly from Baseten's careers page and normalized into a canonical job model.