Liquid AI

Liquid AI

Member of Technical Staff - Edge Inference Engineer

San Francisco · Staff+

Sponsorship not specifiedDetected 178 days ago
RustC++Machine LearningEmbedded Systems

About the role

  • Our Edge Inference team compiles Liquid Foundation Models into optimized machine code that runs on resource-constrained devices: phones, laptops, Raspberry Pis, and watches.

Responsibilities

  • Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability.
  • Implement and optimize inference kernels for CPU, NPU, and GPU architectures across diverse edge hardware
  • Develop quantization strategies (INT4, INT8, FP8) that maximize compression while preserving model quality under strict memory budgets
  • Profile and optimize end-to-end inference pipelines to achieve sub-100ms time-to-first-token on target devices
  • Collaborate with ML researchers to understand model architectures and identify optimization opportunities specific to Liquid Foundation Models

Requirements

  • You can reason about why code is slow before reaching for a profiler.
  • 5+ years of experience in systems programming with strong C++ proficiency
  • Embedded software engineering experience or work on resource-constrained systems
  • Experience with hardware architecture concepts: cache hierarchies, memory bandwidth, SIMD/vectorization

Nice to have

  • Contributions to llama.cpp, ExecuTorch, or similar inference frameworks
  • Experience with Rust for systems programming
  • Background in custom accelerator development (TPU, NPU) or work at companies like SambaNova, Cerebras, Groq, or Google/Amazon accelerator teams
  • Quantitative degree (mathematics, physics, or similar) combined with engineering experience
  • Ship optimizations that achieve measurable latency or memory improvements on at least one target edge device class
  • phones, laptops, Raspberry Pis, and watches.
  • This is high-ownership work where your code ships to production and directly impacts model performance on real devices.

Compensation

  • Competitive base salary with equity in a unicorn-stage company
  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents
  • Financial: 401(k) matching up to 4% of base pay

Benefits

  • Compensation: Competitive base salary with equity in a unicorn-stage company
  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents
  • Time Off: Unlimited PTO plus company-wide Refill Days throughout the year
  • Contribute to llama.cpp and other open-source inference frameworks, including new model architectures (audio, vision)

Company info

  • We partner with enterprises across consumer electronics, automotive, life sciences, and financial services.
  • We are scaling rapidly and need exceptional people to help us get there.
  • We are core contributors to llama.cpp and build the infrastructure that makes efficient on-device AI possible.

This listing is sourced directly from Liquid AI's careers page and normalized into a canonical job model.