River

River

Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

Palo Alto, CA; Austin, TX · Staff+

H1B sponsorship available$200k-$420kDetected 24 days ago
C++Deep LearningLogisticsElectrical Engineering

About the role

  • You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations).

Responsibilities

  • Kernel Generator Development: Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
  • Memory Hierarchy Management: Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
  • HW/SW Co-Design: Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
  • Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
  • Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
  • In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.

Requirements

  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming.
  • Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels.
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM).
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.

Nice to have

  • (We encourage you to apply even if you don't meet all of these)
  • Deep familiarity with implementing microarchitectural optimizations for Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
  • Experience utilizing advanced C++ template metaprogramming or code-generation techniques to automate the creation of heavily parameterized kernel variants.
  • Advanced experience with low-level hardware profiling tools, execution tracing, and utilizing performance counters to identify cache misses, pipeline stalls, and ALU bubbles.
  • Visa Sponsorship: We sponsor visas.
  • We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.

Compensation

  • Depending on background, skills, experience, and location, the expected annual salary range for this position is $200,000 - $420,000 USD.

Benefits

  • River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.
  • Low-Level Compute Optimization: Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware.
  • You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency.

Company info

  • We are scientists, engineers, and builders from the industry's top tech companies and AI labs.
  • We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
  • At River, our mission is to create personal AI owned and shaped by each individual.
  • personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.
  • We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon.

Visa & Work Authorization

  • We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
  • We sponsor visas.

This listing is sourced directly from River's careers page and normalized into a canonical job model.