Friendliai
Software Engineer – AI Inference Engine
San Francisco
Sponsorship not specifiedDetected 129 days ago
PythonC++Machine LearningNLPLLMsAgentic AIElectrical EngineeringResearch
About the role
- In this role, you will focus on designing, implementing, and optimizing GPU kernels and supporting infrastructure for next-generation generative and agentic AI workloads.
- The ideal candidate enjoys pushing performance boundaries and has experience supporting production-scale machine learning applications.
- We power high-throughput, low-latency AI workloads for organizations worldwide and integrate directly with Hugging Face, giving developers instant access to over 500,000 open-source models.
Responsibilities
- Design and optimize custom GPU kernels for AI (e.g., transformer and diffusion) workloads
- Collaborate with cloud and infrastructure engineers to ensure end-to-end inference performance
- Analyze performance bottlenecks across the software and hardware stack, and implement targeted optimizations
- Drive support for new model architectures and tensor compute patterns
- Maintain production-grade performance infrastructure, including profiling, benchmarking, and validation tools
- FriendliAI is building the world's best AI inference platform that makes large language and multi-modal models fast, efficient, and deployable at scale.
- With our world-class inference engine, we are building a platform that the AI industry can actually rely on.
Requirements
- 5+ years of experience in production or high-impact research environments
- Production-level expertise in Python and C++
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
- Experience working with generative AI models such as transformer and diffusion models
- Experience developing machine learning frameworks or performance-critical runtime systems
- Hands-on experience writing and optimizing GPU kernels
- Hands-on experience profiling GPU kernels
Nice to have
- Familiarity with dynamic shape compilation, memory planning, and kernel fusion
- Contributions to inference engines, compilers, or high-performance numerical libraries
- Understanding of multi-GPU and distributed inference strategies
Compensation
- Competitive compensation, startup equity, health insurance, and other benefits.
Benefits
- Flexible working hours
- Health check-up support and top-tier equipment/hardware support
- Competitive compensation, startup equity, health insurance, and other benefits.
Company info
- We are seeking a highly technical Inference Engine Engineer to optimize the performance and efficiency of our core inference engine.
- Your work will directly power the most latency-critical and compute-intensive systems deployed by our customers.
- We are looking for an exceptional engineer with a strong foundation in GPU programming and compiler infrastructure.
- We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology.
Apply directly at Friendliai →Create a free account for alerts like thisView Friendliai immigration profile
This listing is sourced directly from Friendliai's careers page and normalized into a canonical job model.