Unconventionalinc
AI Systems, Model Optimization
Palo Alto, CA I US Remote
Sponsorship not specifiedDetected 15 hours ago
AlgorithmsMachine LearningPyTorchNLPElectrical EngineeringResearchCollaboration
About the role
- As a Member of Technical Staff, AI Systems, Model Optimization, you will develop the path from model architecture to physical silicon.
- You will develop the training techniques, optimization strategies, and infrastructure required to make AI models run efficiently on our novel compute substrates, closing the loop between model design and tapeout.
Responsibilities
- Energy Benchmarking & Performance Modeling: Develop rigorous performance models to evaluate compute, memory, and energy trade-offs. Track pareto-optimality across models and hardware configurations.
- Advanced Mapping & Partitioning: Drive the partitioning and mapping of complex AI models down to hardware.
- Hardware-Aware Training: Develop and apply Quantization-Aware Training (QAT), noise-aware training, and sparsification techniques to adapt models to the physical constraints of our analog compute substrates, including memory footprint, connectivity, precision, and noise.
- GPU Optimization & Kernel Development: Develop and optimize kernels using low-level programming models like CUDA, Triton, or CUTLASS. Profile and debug complex ML codebases to resolve performance bottlenecks (training and inference).
- Energy Benchmarking & Performance Modeling: Develop rigorous performance models to evaluate compute, memory, and energy trade-offs.
- Develop rigorous performance models to evaluate compute, memory, and energy trade-offs.
- Drive the partitioning and mapping of complex AI models down to hardware.
- Develop and apply Quantization-Aware Training (QAT), noise-aware training, and sparsification techniques to adapt models to the physical constraints of our analog compute substrates, including memory footprint, connectivity, precision, and noise.
- Develop and optimize kernels using low-level programming models like CUDA, Triton, or CUTLASS.
Requirements
- Proven experience in profiling, identifying, and resolving performance bottlenecks in complex ML codebases.
- Software Development: Deep experience with PyTorch, including its internals, torch.compile, and distributed data parallel (DDP) / fully sharded data parallel (FSDP) libraries.
- Training Infrastructure: Experience with production-grade training frameworks (e.g., Megatron-LM, DeepSpeed) and distributed training at scale.
- Next-Gen Efficiency: Research or practical experience in advanced approximation/compression techniques beyond standard quantization, including noise-aware or physics-constrained training.
- The Mission: Redefine computing for the next 50 years by solving the fundamental energy limitation of AI at a global scale.
Nice to have
- Deep experience with PyTorch, including its internals, torch.compile, and distributed data parallel (DDP) / fully sharded data parallel (FSDP) libraries.
Benefits
- A comprehensive package including best-in-class health benefits, 401k matching, truly unlimited PTO, and complimentary meals in our Palo Alto office.
Apply directly at Unconventionalinc →Create a free account for alerts like thisView Unconventionalinc immigration profile
This listing is sourced directly from Unconventionalinc's careers page and normalized into a canonical job model.