Nvidia
Senior DevOps Engineer, Cloud Simulation Infrastructure
US, CA, Santa Clara · Senior
Sponsorship not specifiedDetected 5 days ago
PythonCloud PlatformsDockerKubernetesCI/CDDevOpsSite Reliability EngineeringgRPCMachine LearningComputer VisionProduct StrategyRoboticsLoad TestingTest Automation
About the role
- This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF).
- What you'll be doing: Deployment: Deploy full Isaac Sim runtimes within GPU-aware NVCF containers.
- Deploy Runtime Validation: Architect scalable execution layers to conduct runtime behavior-based testing (e.g., drop/grasp tests).
Responsibilities
- Manage container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.
- Deploy Automated Remediation: Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results to ensure quality standards.
- Optimize for performance, addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.
- Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.
- Operational Reliability: Implement atomic update semantics and safe failure handling to ensure validation processes never corrupt the primary asset library.
- Strong ability to design and maintain distributed job lifecycle services (submit/poll/fetch/cancel) and handle asynchronous failure states.
- Experience building "self-healing" or automated remediation workflows.
- Come build the future of autonomous vehicle simulation with us!
- Deploy services to validate USD structure and compliance without runtime overhead.
- Deploy Structural Validation: Deploy services to validate USD structure and compliance without runtime overhead.
Requirements
- What we need to see: BS or MS degree in Computer Science, Computer Engineering, or related field (or equivalent experience).
- Proficiency in Python and systems scripting for test orchestration and pipeline automation.
- Experience with cluster verification frameworks, stress testing, and deployment validation at scale.
- BS or MS degree in Computer Science, Computer Engineering, or related field (or equivalent experience).
Skills
- Scale execution from single-workstation validation to massive, multi-GPU cloud environments.
Compensation
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- The base salary range is 184,000 USD - 287,500 USD.
- Deep familiarity with Isaac Sim, Omniverse, USD, or Sensor RTX workflows.
Benefits
- Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation.
- You will also be eligible for equity and benefits.
Company info
- Ways to stand out from the crowd: Direct experience deploying services on NVCF (NVIDIA Cloud Functions) or DGX Cloud.
This listing is sourced directly from Nvidia's careers page and normalized into a canonical job model.