Etched
Performance Modeling Engineer
San Jose · Staff+
Sponsorship not specifiedDetected 92 days ago
Machine LearningDeep LearningNLPFPGAResearch
About the role
- Our first products are heavily focused on inference.
- Our addressable market is the entirety of inference, unlike many of our competitors.
- We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Responsibilities
- Develop comprehensive performance models and projections for our architecture across varying workloads and configurations
- Drive hardware/software co-optimization by identifying where architectural features can unlock performance improvements
- Run regressions and validate performance models against real systems and silicon
- Inform next-generation architectural decisions by pathfinding across system and silicon options during design, proof-of-concept, and architecting phases
- Strong performance modeling and analysis skills with experience building analytical-based or simulation-based performance models
- Exposure to ASIC, FPGA, or CGRA-based accelerator development and hardware/software co-design principles
- Relocation support for those moving to San Jose (Santana Row)
Requirements
- Deep knowledge of GPU architectures and/or programming models like CUDA
- Experience mapping models to multi-chip inference systems
- Familiarity with transformer model architectures and inference serving optimizations
- Experience with architecture simulators and performance modeling tools (gem5, trace-driven simulators, custom models)
- You may be a good fit if you have at least one of the following:
Benefits
- Medical, dental, and vision packages with generous premium coverage
- $500 per month credit for waiving medical benefits
- Various wellness benefits covering fitness, mental health, and more
- Profile and analyze deep learning workloads on our hardware to identify micro-architectural bottlenecks and influence optimization opportunities
- Experience profiling and analyzing deep learning workloads on hardware accelerators (GPUs, TPUs, ASICs, FPGAs, or others)
Company info
- We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes.
- We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills.
This listing is sourced directly from Etched's careers page and normalized into a canonical job model.