FAL

FAL

Staff Technical Lead for Inference & ML Performance

San Francisco · Staff+

Sponsorship not specifiedDetected 62 days ago
Machine LearningPyTorchResearchLeadership

About the role

  • fal is the generative media ecosystem powering the next generation of AI products.
  • Why this role matters You'll shape the future of fal's inference engine and ensure our generative models achieve best-in-class performance.
  • Personally contribute to critical inference performance enhancements and optimizations.

Responsibilities

  • Set technical direction. Guide your team (kernels, applied performance, ML compilers, distributed inference) to build high-performance inference solutions.
  • Collaborate closely with research & applied ML teams. Influence model inference strategies and deployment techniques.
  • Drive advanced performance optimizations. Implement model parallelism, kernel optimization, and compiler strategies.
  • Guide your team (kernels, applied performance, ML compilers, distributed inference) to build high-performance inference solutions.
  • Collaborate closely with research & applied ML teams.
  • Drive advanced performance optimizations.
  • Implement model parallelism, kernel optimization, and compiler strategies.
  • Lead from the front. You're a respected IC who enjoys getting hands-on with the toughest problems, demonstrating excellence to inspire your team.
  • Experience building inference engines specifically for diffusion and generative media models

Requirements

  • Track record of industry-leading performance improvements (papers, open-source contributions, benchmarks)
  • Leadership experience in scaling technical teams

This listing is sourced directly from FAL's careers page and normalized into a canonical job model.