Nvidia
Senior DL Performance Efficiency Architect
US, CA, Santa Clara · Senior
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Green-card intent detected
Green-card follow-through: 92%
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- This role will bring together model innovation, systems expertise, and hardware awareness to ensure that new capabilities can be delivered within practical constraints of compute, memory, power, and cost.
- The ideal candidate is a hands-on engineer who enjoys finding fundamental bottlenecks, challenging conventional boundaries between disciplines, and turning research ideas into scalable, real-world improvements.
- 5+ years of relevant experience in AI systems, model architecture, computer architecture, high-performance computing, or performance optimization.
Responsibilities
- We are seeking a strong technical leader to drive a unified strategy for making LLMs more efficient from research through deployment.
- You will lead a multidisciplinary effort, establish the technical direction for LLM efficiency, and help shape how future models and computing platforms are designed together.
- Lead cross-layer efforts to improve the efficiency of large language models across model architecture, training and inference systems.
- Analyze how LLM workloads map to GPUs, memory systems, interconnects, and distributed infrastructure, and identify opportunities for model-system-hardware co-design.
- Establish a measurement-driven efficiency roadmap and lead projects from early investigation through production deployment.
- Partner with model researchers, systems engineers, compiler and kernel developers, and hardware architects to influence future model, software, and hardware roadmaps.
- Proven ability to provide technical leadership and drive complex optimization projects from concept to production.
- A first-principles - measure, model, optimize, and deliver - approach to improving LLM efficiency.
Requirements
- MS or PhD degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
- Strong understanding of LLM architectures, training and inference workloads, and the tradeoffs between model quality, computational cost, memory footprint, latency, throughput, and power.
Compensation
- The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
Benefits
- You will also be eligible for equity and benefits.
Company info
- Ways to Stand Out from the Crowd:
This listing is sourced directly from Nvidia's careers page and normalized into a canonical job model.