Chipagents

Chipagents

ML Systems Engineer

San Jose

Sponsorship not specifiedDetected 54 days ago
PythonC++Node.jsMachine LearningPyTorchNLPLLMsAgentic AIA/B TestingSystems EngineeringElectrical EngineeringResearch

About the role

  • This is a technical role focused on low-level systems optimization.

Responsibilities

  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency.
  • Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Requirements

  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
  • Strong proficiency in Python and C++/CUDA
  • hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
  • comfort working at multiple layers of the stack from CUDA kernels to application logic.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Strong systems-level debugging and profiling skills
  • Self-directed problem solver who is interested in working on ambitious optimization challenges.
  • Work on cutting-edge LLM inference optimization problems with real-world production impact.
  • Access to substantial GPU compute resources for experimentation and benchmarking.
  • Collaborate with a world-class team spanning AI research, systems engineering, and EDA.

Nice to have

  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.

Skills

  • The company is a Series A company backed by tier-1 VC firms.
  • ChipAgents is deployed in production to companies that have shipped 16B chips.

Compensation

  • We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.

Benefits

  • $150K/yr - $350K/yr + Offers Equity.
  • Unlimited PTO and full benefits (medical, vision, dental, 401k).
  • Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.

Company info

  • We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform.
  • We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.

This listing is sourced directly from Chipagents's careers page and normalized into a canonical job model.