Designworkstalent

Designworkstalent

AI Training Infrastructure Engineer

Bellevue

Sponsorship not specifiedDetected 19 hours ago
Node.jsDistributed SystemsCloud PlatformsKubernetesMachine LearningPyTorchAgentic AIMLOpsResearch

About the role

  • Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
  • This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

Responsibilities

  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
  • Design and improve systems that increase training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
  • Build tools and automation that improve the developer experience for AI researchers and engineers.
  • Build the infrastructure powering the next generation of AI models and applications.
  • Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

Requirements

  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
  • Strong programming skills and experience working with complex distributed systems.
  • You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.
  • U.S. work authorization is required. Visa sponsorship is not currently available.

Nice to have

  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
  • Experience optimizing GPU utilization, training performance, or distributed system reliability.
  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.
  • Competitive base pay for Bellevue market
  • These awards are allocated based on individual performance
  • Approximately three days per week in the office.
  • U.S. work authorization is required.

Compensation

  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance

Benefits

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience integrating training systems with production machine learning pipelines.
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Visa & Work Authorization

  • U.S. work authorization is required.

This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.