Designworkstalent
AI Training Infrastructure Engineer
Bellevue
Sponsorship not specifiedDetected 19 hours ago
Node.jsDistributed SystemsCloud PlatformsKubernetesMachine LearningPyTorchAgentic AIMLOpsResearch
About the role
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
- This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.
Responsibilities
- Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.
- Design and improve systems that increase training reliability, efficiency, and resource utilization.
- Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
- Build tools and automation that improve the developer experience for AI researchers and engineers.
- Build the infrastructure powering the next generation of AI models and applications.
- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.
Requirements
- Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.
- Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.
- Strong programming skills and experience working with complex distributed systems.
- You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.
- U.S. work authorization is required. Visa sponsorship is not currently available.
Nice to have
- Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
- Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.
- Experience optimizing GPU utilization, training performance, or distributed system reliability.
- Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.
- Competitive base pay for Bellevue market
- These awards are allocated based on individual performance
- Approximately three days per week in the office.
- U.S. work authorization is required.
Compensation
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
Benefits
- Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
- Experience integrating training systems with production machine learning pipelines.
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
Visa & Work Authorization
- U.S. work authorization is required.
Apply directly at Designworkstalent →Create a free account for alerts like thisView Designworkstalent immigration profile
This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.