Designworkstalent
Inference Engineer
Bellevue
Sponsorship not specifiedDetected 19 hours ago
Distributed SystemsCloud PlatformsKubernetesPlatform EngineeringMachine LearningLLMsAgentic AIMLOpsDesign SystemsResearch
About the role
- Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
- This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs.
- You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales.
Responsibilities
- Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.
- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.
- Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
- Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.
- Develop monitoring, alerting, and operational practices to maintain reliable inference services.
- Strong engineering fundamentals and the ability to independently own complex technical problems.
- Build the inference platform powering the next generation of AI applications.
- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.
Requirements
- Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.
- Experience designing reliable distributed systems or production infrastructure.
- U.S. work authorization is required. Visa sponsorship is not currently available.
Nice to have
- Experience with modern inference-serving frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
- Experience optimizing LLM inference workloads or large-scale AI serving platforms.
- Background operating API-based AI products or high-volume production services.
- Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.
- Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning.
- Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization.
- Competitive base pay for Bellevue market
- These awards are allocated based on individual performance
Compensation
- Competitive base pay for Bellevue market
- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance
Benefits
- Experience building and operating production machine learning inference or model-serving systems at scale.
- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.
Visa & Work Authorization
- U.S. work authorization is required.
Apply directly at Designworkstalent →Create a free account for alerts like thisView Designworkstalent immigration profile
This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.