Amazon

Amazon

Senior Inference Engineer, AGI

US, CA, Sunnyvale · Senior

Sponsorship not specified$193k-$262kDetected 19 days ago
Full-Stack DevelopmentAWSDeep LearningNLPLLMsLLMOpsOnboardingCustomer SupportResearchCommunication

Stay score

odds of building a lasting career here

59Sponsors, lottery-bound
Cap-exempt (no lottery)0
Sponsors this role80
Entry-level history0
PERM / green-card track0
Lottery odds (Level IV)94
Fits your clock70

Sponsors, but it's cap-subject — you still face the weighted lottery (~61% per draw at Level IV). Good if you win; have a cap-exempt backup on your list.

Lottery odds assume a STEM candidate.

Personalize to your clock →

Employer immigration record

from this employer's Department of Labor filings

Files H-1B transfers

1 transfer filing in the last year, covering 1 worker. Median labor-condition decision: 24 days. An employer that already files transfers is one that can take over an existing H-1B.

Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.

Community outcomes

No reports yet — be the first to help the next applicant.

About the role

  • Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position.
  • Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
  • The base salary range for this position is listed below.

Requirements

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • Bachelor's degree in computer science or equivalent
  • 2+ years of hands-on experience optimizing inference for neural models - not just using inference frameworks, but profiling and improving them
  • Production track record delivering latency-constrained, real-time inference systems under concurrent load
  • Experience with GPU performance optimization - memory hierarchy, occupancy, KV-cache management, and the accelerator programming model

Nice to have

  • Experience with production LLM/multimodal serving internals (e.g., vLLM, TensorRT-LLM): scheduler, batching, block manager, sampler customization
  • Experience authoring custom GPU kernels (CUTLASS, Triton, raw CUDA/PTX), fused attention (FlashAttention-style), or quantized GEMM
  • Familiarity with model-compression and efficiency techniques - quantization, pruning, distillation, speculative decoding, long-context optimization
  • Experience with distributed training and post-training pipelines (SFT through RL) - parallelism strategies, training stability, and multi-accelerator communication (NCCL, NVLink)
  • Familiarity with multiple hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture choices affect inference latency, memory, and cost
  • Background in speech-to-speech or audio generative models (codec models, autoregressive audio generation), speech recognition, or speech synthesis
  • Experience shipping research to production at scale - models serving real users, not just benchmark results
  • Contributions to open-source inference/kernel projects (vLLM, CUTLASS, FlashAttention, TensorRT-LLM, Triton, or similar)

Compensation

  • The base salary range for this position is listed below.

Company info

  • We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI.

This listing is sourced directly from Amazon's careers page and normalized into a canonical job model.