Inferact

Inferact

Member of Technical Staff, Developer Relations

San Francisco · Staff+

H1B sponsorship available$200k-$400kDetected 34 days ago
Distributed SystemsMachine LearningPyTorchLLMsMLOpsLogistics

About the role

  • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.

Requirements

  • Experience with vLLM or adjacent inference technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms, or similar systems.
  • Ability to write and teach for practitioners without sounding like a content marketer.

Nice to have

  • Prior work in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
  • Experience creating technical content that teaches reusable mental models, not just product features.
  • Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.
  • Written widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.
  • Built demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.
  • Created practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.
  • Built a durable personal portfolio that demonstrates technical depth, taste, and a strong point of view.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.

Skills

  • Skills and Qualifications

Compensation

  • Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Benefits

  • Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Company info

  • Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
  • We're looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference.

Visa & Work Authorization

  • We sponsor visas on a case-by-case basis.

This listing is sourced directly from Inferact's careers page and normalized into a canonical job model.