Flexai

Flexai

Staff AI Runtime Engineer

Santa Clara, CA · Staff+

Sponsorship not specifiedDetected 119 days ago
PythonGoRustC++Node.jsDistributed SystemsKubernetesCI/CDDeep LearningTensorFlowPyTorchLLMsResearchLeadership

About the role

  • This is a hands-on leadership role - perfect for a systems-minded software engineer who thrives at the intersection of AI workloads, runtimes, and performance-critical infrastructure.

Responsibilities

  • Lead Runtime Design & Development: Own the core runtime architecture supporting AI training and inference at scale.
  • Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
  • Optimize distributed training reliability, orchestration, and job-level fault tolerance.
  • Drive Performance at Scale: Profile and enhance low-level system performance across training and inference pipelines.
  • Build Internal Tooling & Frameworks: Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
  • Collaborate & Mentor: Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.

Nice to have

  • Contributions to PyTorch internals or open-source DL infrastructure projects.
  • Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
  • Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
  • Background in systems research, compilers, or runtime architecture for HPC or ML.
  • Start up previous experience This position is In-Person and located at our Santa Clara, CA Office.
  • Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
  • Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team's capabilities.
  • What You'll Need to Be Successful 8+ years of experience in systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.

Skills

  • It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.
  • Our teams are strategically distributed across
  • to deliver more compute with less complexity.
  • Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.

Benefits

  • A competitive salary and benefits package Work on cutting-edge AI infrastructure Build products used by developers and enterprises High ownership, fast execution, real impact Collaborative, high-caliber team
  • Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
  • Proven experience optimizing and scaling deep learning runtimes (e.g. PyTorch, TensorFlow, JAX) for large-scale training and/or inference.

This listing is sourced directly from Flexai's careers page and normalized into a canonical job model.