Inferact

Inferact

Member of Technical Staff, Kernel Engineering

San Francisco · Staff+

H1B sponsorship available$200k-$400kDetected 181 days ago
PythonC++Machine LearningLogistics

About the role

  • We're looking for a performance engineer to squeeze every FLOP out of modern accelerators.
  • You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world.
  • Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon.

Responsibilities

  • When hardware vendors develop new chips, they integrate with vLLM.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
  • Proficiency in C++ and Python with demonstrated ability to write high-performance code.
  • Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.

Nice to have

  • Experience with ML-specific kernel optimization (FlashAttention, fused kernels).
  • Knowledge of quantization techniques (INT8, FP8, mixed-precision).
  • Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).
  • Experience with compiler technologies (LLVM, MLIR, XLA).
  • Kernel-related contributions to vLLM or other inference engine projects.
  • Contributions to open-source GPU, ML systems, or compiler optimization projects
  • Written deep technical blogs on GPU optimization.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.

Compensation

  • Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Benefits

  • Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Company info

  • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
  • Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

Visa & Work Authorization

  • We sponsor visas on a case-by-case basis.

This listing is sourced directly from Inferact's careers page and normalized into a canonical job model.