Nvidia
Senior Systems Software Engineer - Fleet Debuggability
US, CA, Santa Clara · Senior
Sponsorship not specifiedDetected 2 days ago
PythonRustC++Code ReviewGitElasticsearchLinuxPrometheusGrafanaDeep LearningLLMsProject ManagementJiraCommunicationCollaboration
About the role
- NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing.
- More recently, GPU deep learning ignited modern deep learning - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world.
- The solution should normalize, correlate, and reason over logs spanning multiple components, trays, or racks including NVIDIA's GPUs, CPUs, Network products.
Responsibilities
- You will design, architect, and build infrastructure, tooling, analytics on how to collect multi-rack scale logs.
- Architect, Design, build, fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
- Develop tooling to collect, normalize, and time-align logs from heterogeneous sources - kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs - over both in-band and out-of-band channels.
- Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, so triage is repeatable rather than tribal knowledge.
- Develop debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.
- Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
- Partner with all matrixed organizations - developers, SWQA, and product engineering - in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
- Experience standing up follow-the-sun support organizations with measurable response SLAs Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code.
Requirements
- 10+ years in the software industry with specialization in system software and/or firmware development.
- Proven track record of shipping scalable server products or fleet-wide experience.
- Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira.
- Hands-on experience with out-of-band management and platform interfaces - BMC, Redfish, IPMI, SEL - and an understanding of in-band vs. out-of-band trade-offs.
- Hands-on experience with x86/ARM system architecture and coding (C/C++, Python).
- Experience with SCM (Git, Perforce) and project management tools (Jira).
Skills
- keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
Compensation
- The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.
Benefits
- You will also be eligible for equity and benefits.
Company info
- Today, we are increasingly known as "the AI computing company." We're looking to grow our company and establish teams with the most thoughtful people in the world.
- We are the Datacenter System Software team, and we are looking for a highly motivated, creative Senior Engineer o drive Fleet Scale Debuggability end to end.
- The logs shall be fetched inband or out of band and should help triage fleet level issues seen by our customers.
Equal opportunity
- equal opportunity employer.
This listing is sourced directly from Nvidia's careers page and normalized into a canonical job model.