Chipagents
ML Systems Engineer
San Jose
Sponsorship not specifiedDetected 54 days ago
PythonC++Node.jsMachine LearningPyTorchNLPLLMsAgentic AIA/B TestingSystems EngineeringElectrical EngineeringResearch
About the role
- This is a technical role focused on low-level systems optimization.
Responsibilities
- Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
- Implement and benchmark concrete inference optimizations.
- Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
- Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
- You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency.
- Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.
Requirements
- B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
- Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
- Strong proficiency in Python and C++/CUDA
- hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
- Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
- comfort working at multiple layers of the stack from CUDA kernels to application logic.
- Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
- Strong systems-level debugging and profiling skills
- Self-directed problem solver who is interested in working on ambitious optimization challenges.
- Work on cutting-edge LLM inference optimization problems with real-world production impact.
- Access to substantial GPU compute resources for experimentation and benchmarking.
- Collaborate with a world-class team spanning AI research, systems engineering, and EDA.
Nice to have
- Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
Skills
- The company is a Series A company backed by tier-1 VC firms.
- ChipAgents is deployed in production to companies that have shipped 16B chips.
Compensation
- We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.
Benefits
- $150K/yr - $350K/yr + Offers Equity.
- Unlimited PTO and full benefits (medical, vision, dental, 401k).
- Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.
Company info
- We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform.
- We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.
Apply directly at Chipagents →Create a free account for alerts like thisView Chipagents immigration profile
This listing is sourced directly from Chipagents's careers page and normalized into a canonical job model.