Inferact

Inferact

Member of Technical Staff, TPU Performance Engineering

San Francisco · Staff+

H1B sponsorship available$200k-$400kDetected 34 days ago
Machine LearningPyTorchLLMsLogistics

About the role

  • We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs.
  • You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks.

Responsibilities

  • You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.
  • Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.

Requirements

  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.
  • Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.

Nice to have

  • Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.
  • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
  • Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.
  • Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.
  • Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.
  • Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
  • Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.

Skills

  • Skills and Qualifications

Compensation

  • Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Benefits

  • Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Company info

  • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
  • Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

Visa & Work Authorization

  • We sponsor visas on a case-by-case basis.

This listing is sourced directly from Inferact's careers page and normalized into a canonical job model.