Inferact
Member of Technical Staff, TPU Performance Engineering
San Francisco · Staff+
H1B sponsorship available$200k-$400kDetected 34 days ago
Machine LearningPyTorchLLMsLogistics
About the role
- We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs.
- You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks.
Responsibilities
- You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware.
- Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable.
Requirements
- Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths.
- Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work.
Nice to have
- Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems.
- Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
- Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation.
- Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs.
- Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure.
- Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
- Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
- Visa sponsorship: We sponsor visas on a case-by-case basis.
Skills
- Skills and Qualifications
Compensation
- Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Benefits
- Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Company info
- Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
- Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
Visa & Work Authorization
- We sponsor visas on a case-by-case basis.
Apply directly at Inferact →Create a free account for alerts like thisView Inferact immigration profile
This listing is sourced directly from Inferact's careers page and normalized into a canonical job model.