Pika

Pika

ML Engineer, Inference & Optimization

Palo Alto HQ · Senior

Sponsorship not specifiedDetected 28 days ago
Code ReviewMachine LearningDeep LearningLLMsResearchCommunicationCollaboration

About the role

  • We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products.
  • In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies.
  • Your efforts will play a foundational role in powering the next generation of Pika's video and language models.

Responsibilities

  • Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
  • Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
  • Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
  • Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
  • Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
  • able to drive shared goals across research and engineering functions.
  • Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
  • A supportive and collaborative office culture-we're all building and launching together

Requirements

  • GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
  • Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
  • Familiarity with video generation (videogen) models and large language models (LLMs).

Nice to have

  • Experience with high-throughput video or real-time streaming model deployment
  • Familiarity with distributed training and optimization toolkits
  • Startup or rapid prototyping experience
  • We're a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.
  • Experience in enhancing training efficiency, stability, or resource optimization for large models.

Compensation

  • Competitive salary in the AI industry

Benefits

  • Equity in a fast-growing startup shaping the future of AI
  • Comprehensive health benefits, monthly stipends, company retreats
  • (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
  • Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.

Company info

  • At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible.
  • Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI.
  • We work from our Palo Alto office 3-5 days a week and welcome applicants who are eager to contribute onsite.

This listing is sourced directly from Pika's careers page and normalized into a canonical job model.