River
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)
Palo Alto, CA; Austin, TX · Staff+
H1B sponsorship available$200k-$420kDetected 24 days ago
C++Deep LearningLogisticsElectrical Engineering
About the role
- You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations).
Responsibilities
- Kernel Generator Development: Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
- Memory Hierarchy Management: Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
- HW/SW Co-Design: Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
- Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
- Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
- In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.
Requirements
- Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming.
- Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels.
- Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM).
- Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.
Nice to have
- (We encourage you to apply even if you don't meet all of these)
- Deep familiarity with implementing microarchitectural optimizations for Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
- Experience utilizing advanced C++ template metaprogramming or code-generation techniques to automate the creation of heavily parameterized kernel variants.
- Advanced experience with low-level hardware profiling tools, execution tracing, and utilizing performance counters to identify cache misses, pipeline stalls, and ALU bubbles.
- Visa Sponsorship: We sponsor visas.
- We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
Compensation
- Depending on background, skills, experience, and location, the expected annual salary range for this position is $200,000 - $420,000 USD.
Benefits
- River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.
- Low-Level Compute Optimization: Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware.
- You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency.
Company info
- We are scientists, engineers, and builders from the industry's top tech companies and AI labs.
- We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
- At River, our mission is to create personal AI owned and shaped by each individual.
- personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.
- We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon.
Visa & Work Authorization
- We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
- We sponsor visas.
This listing is sourced directly from River's careers page and normalized into a canonical job model.