FAL
Staff Technical Lead for Inference & ML Performance
San Francisco · Staff+
Sponsorship not specifiedDetected 62 days ago
Machine LearningPyTorchResearchLeadership
About the role
- fal is the generative media ecosystem powering the next generation of AI products.
- Why this role matters You'll shape the future of fal's inference engine and ensure our generative models achieve best-in-class performance.
- Personally contribute to critical inference performance enhancements and optimizations.
Responsibilities
- Set technical direction. Guide your team (kernels, applied performance, ML compilers, distributed inference) to build high-performance inference solutions.
- Collaborate closely with research & applied ML teams. Influence model inference strategies and deployment techniques.
- Drive advanced performance optimizations. Implement model parallelism, kernel optimization, and compiler strategies.
- Guide your team (kernels, applied performance, ML compilers, distributed inference) to build high-performance inference solutions.
- Collaborate closely with research & applied ML teams.
- Drive advanced performance optimizations.
- Implement model parallelism, kernel optimization, and compiler strategies.
- Lead from the front. You're a respected IC who enjoys getting hands-on with the toughest problems, demonstrating excellence to inspire your team.
- Experience building inference engines specifically for diffusion and generative media models
Requirements
- Track record of industry-leading performance improvements (papers, open-source contributions, benchmarks)
- Leadership experience in scaling technical teams
This listing is sourced directly from FAL's careers page and normalized into a canonical job model.