Friendliai

Friendliai

Software Engineer – AI Inference Engine

San Francisco

Sponsorship not specifiedDetected 129 days ago
PythonC++Machine LearningNLPLLMsAgentic AIElectrical EngineeringResearch

About the role

  • In this role, you will focus on designing, implementing, and optimizing GPU kernels and supporting infrastructure for next-generation generative and agentic AI workloads.
  • The ideal candidate enjoys pushing performance boundaries and has experience supporting production-scale machine learning applications.
  • We power high-throughput, low-latency AI workloads for organizations worldwide and integrate directly with Hugging Face, giving developers instant access to over 500,000 open-source models.

Responsibilities

  • Design and optimize custom GPU kernels for AI (e.g., transformer and diffusion) workloads
  • Collaborate with cloud and infrastructure engineers to ensure end-to-end inference performance
  • Analyze performance bottlenecks across the software and hardware stack, and implement targeted optimizations
  • Drive support for new model architectures and tensor compute patterns
  • Maintain production-grade performance infrastructure, including profiling, benchmarking, and validation tools
  • FriendliAI is building the world's best AI inference platform that makes large language and multi-modal models fast, efficient, and deployable at scale.
  • With our world-class inference engine, we are building a platform that the AI industry can actually rely on.

Requirements

  • 5+ years of experience in production or high-impact research environments
  • Production-level expertise in Python and C++
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
  • Experience working with generative AI models such as transformer and diffusion models
  • Experience developing machine learning frameworks or performance-critical runtime systems
  • Hands-on experience writing and optimizing GPU kernels
  • Hands-on experience profiling GPU kernels

Nice to have

  • Familiarity with dynamic shape compilation, memory planning, and kernel fusion
  • Contributions to inference engines, compilers, or high-performance numerical libraries
  • Understanding of multi-GPU and distributed inference strategies

Compensation

  • Competitive compensation, startup equity, health insurance, and other benefits.

Benefits

  • Flexible working hours
  • Health check-up support and top-tier equipment/hardware support
  • Competitive compensation, startup equity, health insurance, and other benefits.

Company info

  • We are seeking a highly technical Inference Engine Engineer to optimize the performance and efficiency of our core inference engine.
  • Your work will directly power the most latency-critical and compute-intensive systems deployed by our customers.
  • We are looking for an exceptional engineer with a strong foundation in GPU programming and compiler infrastructure.
  • We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology.

This listing is sourced directly from Friendliai's careers page and normalized into a canonical job model.