Plaud.ai

Plaud.ai

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

San Francisco, CA · Full-time

Sponsorship not specified$170k-$320kDetected 75 days ago
Node.jsDistributed SystemsAlgorithmsWebSocketsMachine LearningLLMsComplianceVoIPPro ToolsHIPAACollaboration

About the role

  • Plaud Inc. is a Delaware-incorporated, San Francisco-based company pushing the boundary of human-AI intelligence through a hardware-software combination.
  • With full ISO 27001, ISO 27701, SOC 2, GDPR, EN18031, and HIPAA compliances, Plaud is committed to the highest standards of data security and privacy protection.
  • Plaud is and will continue to be an equal opportunity employer.

Requirements

  • Have practical experience with continuous batching, KV cache management (e.g., PagedAttention), and stateful connections necessary for real-time conversational AI.
  • Hands-on experience with post-training quantization (PTQ), deploying models in FP8, INT8, AWQ, or GPTQ, without degrading audio naturalness or ASR accuracy.

Compensation

  • $170K - $320K base salary + performance bonus + Equity.
  • Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.

Benefits

  • Founding Team Initiative: Opportunity to be an early, foundational member of our core SpeechLLM lab, with meaningful ownership and impact on a fast-growing startup.
  • Competitive Compensation: $170K - $320K base salary + performance bonus + Equity.
  • Comprehensive Benefits: Top-tier healthcare for employees and dependents, including dental and vision, and a generous employer subsidy.
  • Retirement Planning: 401(k) plan for full-time employees with company matching.
  • Paid Time Off: Unlimited PTO, plus 13 paid holidays.
  • New Parent Leave: 12 weeks of paid time off to spend time with your new family, regardless of gender.
  • Gear & Perks: Choice of top-of-the-line laptops/workstations, annual offsites, and a fully stocked office.
  • We do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristics.
  • Deep, under-the-hood familiarity with modern LLM serving frameworks like vLLM, TensorRT-LLM, SGLang, or NVIDIA Triton Inference Server (bonus points for active open-source contributions to these repositories).

Equal opportunity

  • equal opportunity employer.

This listing is sourced directly from Plaud.ai's careers page and normalized into a canonical job model.