Designworkstalent

Designworkstalent

Inference Engineer

Bellevue

Sponsorship not specifiedDetected 19 hours ago
Distributed SystemsCloud PlatformsKubernetesPlatform EngineeringMachine LearningLLMsAgentic AIMLOpsDesign SystemsResearch

About the role

  • Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
  • This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs.
  • You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales.

Responsibilities

  • Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.
  • Design systems that maximize GPU utilization while maintaining predictable performance and reliability.
  • Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.
  • Develop monitoring, alerting, and operational practices to maintain reliable inference services.
  • Strong engineering fundamentals and the ability to independently own complex technical problems.
  • Build the inference platform powering the next generation of AI applications.
  • Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

Requirements

  • Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.
  • Experience designing reliable distributed systems or production infrastructure.
  • U.S. work authorization is required. Visa sponsorship is not currently available.

Nice to have

  • Experience with modern inference-serving frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
  • Experience optimizing LLM inference workloads or large-scale AI serving platforms.
  • Background operating API-based AI products or high-volume production services.
  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.
  • Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning.
  • Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization.
  • Competitive base pay for Bellevue market
  • These awards are allocated based on individual performance

Compensation

  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance

Benefits

  • Experience building and operating production machine learning inference or model-serving systems at scale.
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Visa & Work Authorization

  • U.S. work authorization is required.

This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.