Liquid AI
Member of Technical Staff - Edge Inference Engineer
San Francisco · Staff+
Sponsorship not specifiedDetected 178 days ago
RustC++Machine LearningEmbedded Systems
About the role
- Our Edge Inference team compiles Liquid Foundation Models into optimized machine code that runs on resource-constrained devices: phones, laptops, Raspberry Pis, and watches.
Responsibilities
- Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability.
- Implement and optimize inference kernels for CPU, NPU, and GPU architectures across diverse edge hardware
- Develop quantization strategies (INT4, INT8, FP8) that maximize compression while preserving model quality under strict memory budgets
- Profile and optimize end-to-end inference pipelines to achieve sub-100ms time-to-first-token on target devices
- Collaborate with ML researchers to understand model architectures and identify optimization opportunities specific to Liquid Foundation Models
Requirements
- You can reason about why code is slow before reaching for a profiler.
- 5+ years of experience in systems programming with strong C++ proficiency
- Embedded software engineering experience or work on resource-constrained systems
- Experience with hardware architecture concepts: cache hierarchies, memory bandwidth, SIMD/vectorization
Nice to have
- Contributions to llama.cpp, ExecuTorch, or similar inference frameworks
- Experience with Rust for systems programming
- Background in custom accelerator development (TPU, NPU) or work at companies like SambaNova, Cerebras, Groq, or Google/Amazon accelerator teams
- Quantitative degree (mathematics, physics, or similar) combined with engineering experience
- Ship optimizations that achieve measurable latency or memory improvements on at least one target edge device class
- phones, laptops, Raspberry Pis, and watches.
- This is high-ownership work where your code ships to production and directly impacts model performance on real devices.
Compensation
- Competitive base salary with equity in a unicorn-stage company
- Health: We pay 100% of medical, dental, and vision premiums for employees and dependents
- Financial: 401(k) matching up to 4% of base pay
Benefits
- Compensation: Competitive base salary with equity in a unicorn-stage company
- Health: We pay 100% of medical, dental, and vision premiums for employees and dependents
- Time Off: Unlimited PTO plus company-wide Refill Days throughout the year
- Contribute to llama.cpp and other open-source inference frameworks, including new model architectures (audio, vision)
Company info
- We partner with enterprises across consumer electronics, automotive, life sciences, and financial services.
- We are scaling rapidly and need exceptional people to help us get there.
- We are core contributors to llama.cpp and build the infrastructure that makes efficient on-device AI possible.
Apply directly at Liquid AI →Create a free account for alerts like thisView Liquid AI immigration profile
This listing is sourced directly from Liquid AI's careers page and normalized into a canonical job model.