Amazon
Senior Inference Engineer, AGI
US, CA, Sunnyvale · Senior
Stay score
odds of building a lasting career here
Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position.
- Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
- The base salary range for this position is listed below.
Requirements
- PhD, or Master's degree and 6+ years of applied research experience
- Experience programming in Java, C++, Python or related language
- 2+ years of hands-on experience optimizing inference for neural models - not just using inference frameworks, but profiling and improving them
- Production track record delivering latency-constrained, real-time inference systems under concurrent load
- Experience with GPU performance optimization - memory hierarchy, occupancy, KV-cache management, and the accelerator programming model
Nice to have
- Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
- Experience with production LLM/multimodal serving internals (e.g., vLLM, TensorRT-LLM): scheduler, batching, block manager, sampler customization
- Experience authoring custom GPU kernels (CUTLASS, Triton, raw CUDA/PTX), fused attention (FlashAttention-style), or quantized GEMM
- Familiarity with model-compression and efficiency techniques - quantization, pruning, distillation, speculative decoding, long-context optimization
- Experience with distributed training and post-training pipelines (SFT through RL) - parallelism strategies, training stability, and multi-accelerator communication (NCCL, NVLink)
- Familiarity with multiple hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture choices affect inference latency, memory, and cost
- Background in speech-to-speech or audio generative models (codec models, autoregressive audio generation), speech recognition, or speech synthesis
- Experience shipping research to production at scale - models serving real users, not just benchmark results
Compensation
- The base salary range for this position is listed below.
Company info
- We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI.
This listing is sourced directly from Amazon's careers page and normalized into a canonical job model.