Nvidia
Senior Site Reliability Engineer - Storage
US, CA, Santa Clara · Senior
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Green-card intent detected
Green-card follow-through: 92%
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization.
- The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
- Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars.
Responsibilities
- Design, implement an on-prem HPC infrastructure supplemented with cloud computing to support the growing IT needs of NVIDIA.
- Design and implement scalable and efficient Storage solutions tailored for data-intensive applications, optimizing performance and cost-effectiveness.
- Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.
- Document the general procedures and practices, perform technology evaluations, related to distributed file systems.
- Collaborate across teams to better understand developers' workflows and gather their infrastructure requirements.
- Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization.
- Design, deployment and management of Enterprise NAS solutions like NetApp, Pure Storage and S3 based storage such as cloudian MinIO.
- You will be responsible for crafting and deploying distributed storage solutions, build automation tools, and ensuring the efficient operations of our growing IT ecosystem.
- You will collaborate closely with engineering teams to align infrastructure with their evolving needs, document best practices, and contribute to the success of ground breaking projects.
Requirements
- What we need see: BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years of experience or Ph.D. with 3 years of experience.
- 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.
- Experience with one or more parallel or distributed filesystems such as Lustre, GPFS.
- Experience with multiple monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix, etc.
- Prior Experience with HPC cluster management tools such as Slurm, PBS, LSF, etc.
- Experience with containerization technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).
- BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years of experience or Ph.D. with 3 years of experience.
Compensation
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
Company info
- Ways To Stand Out Of The Crowd: Background with RDMA (InfiniBand or RoCE) fabrics.
Equal opportunity
- NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer.
This listing is sourced directly from Nvidia's careers page and normalized into a canonical job model.