Plaud.ai
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, CA · Full-time
Sponsorship not specified$170k-$320kDetected 75 days ago
Node.jsDistributed SystemsAlgorithmsWebSocketsMachine LearningLLMsComplianceVoIPPro ToolsHIPAACollaboration
About the role
- Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human-AI intelligence through a hardware-software combination.
- With full ISO 27001, ISO 27701, SOC 2, GDPR, EN18031, and HIPAA compliances, Plaud is committed to the highest standards of data security and privacy protection.
- Plaud is and will continue to be an equal opportunity employer.
Requirements
- Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.
- Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.
Compensation
- $170K - $320K base salary + performance bonus + Equity.
- Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.
Benefits
- Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup.
- Competitive Compensation: $170K - $320K base salary + performance bonus + Equity.
- Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy.
- Retirement Planning: 401(k) plan for full-time employees with company matching.
- Paid Time Off: Unlimited PTO, plus 13 paid holidays.
- New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender.
- Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.
- We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.
- Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).
Equal opportunity
- equal opportunity employer.
Apply directly at Plaud.ai →Create a free account for alerts like thisView Plaud.ai immigration profile
This listing is sourced directly from Plaud.ai's careers page and normalized into a canonical job model.