Designworkstalent

Designworkstalent

GPU Performance / Kernel Engineer

Bellevue

Sponsorship not specifiedDetected 18 hours ago
Cloud PlatformsPlatform EngineeringMachine LearningAgentic AISystems Engineering

About the role

  • Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization.
  • This role focuses on improving GPU utilization, reducing latency, and maximizing throughput across training and inference environments by tuning kernels, identifying performance bottlenecks, and driving efficiency across the GPU fleet.
  • This is a high-impact engineering role focused on extracting maximum performance from large-scale GPU infrastructure.

Responsibilities

  • Profile, analyze, and optimize GPU kernels to improve latency, throughput, and overall utilization.
  • Develop benchmarking methodologies and performance measurement practices across GPU infrastructure.
  • Ability to independently own technically complex problems and drive solutions in a fast-moving engineering environment.
  • Optimize the performance layer behind one of the industry's most advanced AI infrastructure platforms.
  • Collaborate with world-class engineers building the infrastructure powering the next generation of AI applications.

Requirements

  • Strong experience with GPU kernel development and performance optimization using technologies such as CUDA, ROCm, or comparable GPU programming frameworks.
  • Demonstrated experience improving GPU utilization, reducing latency, or increasing throughput for production AI workloads.
  • Strong understanding of GPU architecture, memory hierarchy, parallel computing, and the data path from application layer to hardware execution.
  • Experience profiling and debugging performance issues in complex AI or distributed computing environments.
  • U.S. work authorization is required. Visa sponsorship is not currently available.

Nice to have

  • Experience optimizing workloads across multiple GPU platforms, including NVIDIA and AMD architectures.
  • Experience with GPU compiler technologies, runtime optimization, or low-level systems performance.
  • Contributions to open-source GPU performance projects, compiler tooling, or AI systems optimization.
  • Background working with large-scale AI training, inference platforms, HPC environments, or cloud GPU infrastructure.
  • Familiarity with GPU profiling and optimization tools such as Nsight Systems, Nsight Compute, ROCm profiling tools, or similar technologies.
  • Competitive base pay for Bellevue market
  • These awards are allocated based on individual performance
  • Approximately three days per week in the office.

Compensation

  • Competitive base pay for Bellevue market
  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and stock. These awards are allocated based on individual performance

Benefits

  • Work closely with AI infrastructure, machine learning, and platform engineering teams to understand workload characteristics and optimize system behavior.
  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Company info

  • Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

Visa & Work Authorization

  • U.S. work authorization is required.

This listing is sourced directly from Designworkstalent's careers page and normalized into a canonical job model.