Genmo
GPU Performance Engineer
San Francisco HQ
Sponsorship not specifiedDetected 370 days ago
PythonC++SassMachine LearningNLPElectrical Engineering
About the role
- You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups.
- This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.
Responsibilities
- Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation
- Optimize cold start latency from seconds to milliseconds for our serving infrastructure
- Collaborate with ML engineers to optimize model implementations
- Implement custom memory pooling and allocation strategies
- Share optimization techniques and build performance culture across teams
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
- 5+ years systems programming experience with 3+ years focused on GPU optimization
- Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)
- Track record of achieving significant performance improvements (5-10x)
- Experience with Python and C++ in production environments
- Experience with Triton kernel development
- Knowledge of CUTLASS or similar high-performance libraries
- Strong CUDA programming skills with production kernel development
- Deep understanding of GPU architecture (memory hierarchy, SMs, warps)
- Background in ML-specific optimizations (attention, transformers)
- RDMA/InfiniBand optimization experience
- Contributions to GPU libraries or frameworks
- Low-level debugging skills (PTX/SASS reading)
This listing is sourced directly from Genmo's careers page and normalized into a canonical job model.