Orion Innovation

Orion Innovation

Senior ML Infrastructure Engineer

Edison, NJ · Senior

Sponsorship not specifiedDetected 22 days ago
PythonNode.jsAzureDockerKubernetesTerraformHelmPlatform EngineeringMachine LearningPyTorchNLP

About the role

  • The platform runs on Azure Kubernetes Service (AKS) with dedicated GPU node pools, uses KEDA for event-driven autoscaling, and integrates with Azure Data Lake Storage Gen2 and Azure OpenAI.
  • This is a hands-on role that sits at the intersection of platform engineering and applied ML, and requires someone who is equally comfortable debugging a CUDA out-of-memory error and designing a Kubernetes autoscaling policy.

Responsibilities

  • Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and engineering, with a unique combination of agility, scale, and maturity.
  • Multi-node-pool design, taint/toleration, autoscaler, GPU node pools (NC/ND series)
  • As the Senior ML Infrastructure Engineer the resource will own the end-to-end infrastructure layer - from GPU cluster configuration and CUDA runtime management to Kubernetes job orchestration and model serving.

Requirements

  • raw kernel dev not required, PyTorch (GPU inference)

Nice to have

  • What information we collect during our application and recruitment process and why we collect it;
  • How we handle that information
  • How to access and update that information.

Visa & Work Authorization

  • race, color, creed, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, citizenship status, disability status, genetic information, protected veteran status, or any other ch

This listing is sourced directly from Orion Innovation's careers page and normalized into a canonical job model.