Orion Innovation
Senior ML Infrastructure Engineer
Edison, NJ · Senior
Sponsorship not specifiedDetected 22 days ago
PythonNode.jsAzureDockerKubernetesTerraformHelmPlatform EngineeringMachine LearningPyTorchNLP
About the role
- The platform runs on Azure Kubernetes Service (AKS) with dedicated GPU node pools, uses KEDA for event-driven autoscaling, and integrates with Azure Data Lake Storage Gen2 and Azure OpenAI.
- This is a hands-on role that sits at the intersection of platform engineering and applied ML, and requires someone who is equally comfortable debugging a CUDA out-of-memory error and designing a Kubernetes autoscaling policy.
Responsibilities
- Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and engineering, with a unique combination of agility, scale, and maturity.
- Multi-node-pool design, taint/toleration, autoscaler, GPU node pools (NC/ND series)
- As the Senior ML Infrastructure Engineer the resource will own the end-to-end infrastructure layer - from GPU cluster configuration and CUDA runtime management to Kubernetes job orchestration and model serving.
Requirements
- raw kernel dev not required, PyTorch (GPU inference)
Nice to have
- What information we collect during our application and recruitment process and why we collect it;
- How we handle that information
- How to access and update that information.
Visa & Work Authorization
- race, color, creed, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, citizenship status, disability status, genetic information, protected veteran status, or any other ch
Apply directly at Orion Innovation →Create a free account for alerts like thisView Orion Innovation immigration profile
This listing is sourced directly from Orion Innovation's careers page and normalized into a canonical job model.