Zyphra

Zyphra

Research Engineer - AI Performance & Kernel Optimization

San Francisco

Sponsorship not specified$10k-$20kDetected 127 days ago
AWSMachine LearningNLPElectrical EngineeringResearchExperimental DesignCommunicationCollaboration

About the role

  • This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.
  • Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs
  • Performance tuning for training and inference stacks across GPUs and other accelerators

Responsibilities

  • Relocation and immigration support on a case-by-case basis

Requirements

  • Excellent low-level performance intuition and the ability to reason about hardware-software interactions
  • Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams
  • Strong understanding of distributed training systems and parallelism schemes, including data parallelism, tensor/model parallelism, pipeline parallelism, sharding, and communication/computation overlap
  • Experience with performance engineering in other demanding parallel computing environments such as HPC, quantitative finance, scientific computing, graphics, compilers, or numerical simulation
  • Familiarity with infrastructure underlying large-scale training and inference, including collective communication libraries, and runtime performance analysis
  • Strong systems intuition around memory hierarchy, bandwidth constraints, kernel fusion, launch overhead, communication overhead, and hardware utilization
  • Experience using profiling and debugging tools to drive performance improvements
  • Background in a highly technical field such as physics, mathematics, theoretical computer science, computer science, or electrical engineering
  • Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued

Skills

  • Experience writing highly performant GPU kernels at any level of abstraction-PTX, CUDA, HIP, Triton, or other kernel DSLs
  • Experience optimizing ML workloads for large-scale training, ideally in language model pretraining or inference environments
  • Experience with non-NVIDIA accelerator hardware, such as AMD, AWS Trainium, Google TPU, Qualcomm, ARM, Intel, and custom ASICs
  • Any HPC experience is a strong plus
  • We strongly value new and crazy ideas and are very willing to bet big on new ideas
  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible
  • We all enjoy what we do and love discussing AI

Compensation

  • Competitive compensation and 401(k) plan

Benefits

  • Comprehensive medical, dental, vision, and FSA plans
  • Unlimited PTO and company holidays

Company info

  • ZYPHRA IS AN ARTIFICIAL INTELLIGENCE COMPANY BASED IN SAN FRANCISCO, CALIFORNIA.

This listing is sourced directly from Zyphra's careers page and normalized into a canonical job model.