Material World
Inference Performance Engineer
New York, NY
Sponsorship not specifiedDetected 69 days ago
PythonGoRustC++NLPLLMs
About the role
- Serving frontier models at scale requires solving novel systems problems at every layer of the stack.
Responsibilities
- Build and improve the inference runtime
- Design scheduling, continuous batching, KV cache, and prefill/decode disaggregation
- Implement low-precision kernels and speculative decoding
- Drive throughput, latency, and cost per token
- Collaborate with hardware teams on kernels, operators, and graph optimizations
- Own the OpenAI-compatible API surface and serving protocol
- Build benchmarking, profiling, and regression infrastructure
- Relocation support
Requirements
- Software engineering experience: Rust, Go, Python, or C++
- Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp, or custom runtimes
- Experience with low-precision inference (FP8, FP4, INT4)
Compensation
- Top-tier compensation structured to recognize and retain the best talent
Benefits
- Meaningful equity
- Comprehensive medical, dental, vision, life, and disability insurance
- Parental leave for all new parents, including adoptive and surrogate journeys
Equal opportunity
- We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.
Apply directly at Material World →Create a free account for alerts like thisView Material World immigration profile
This listing is sourced directly from Material World's careers page and normalized into a canonical job model.