Netpreme

Netpreme

Member of Technical Staff, Performance Modeling

Santa Clara, CA or Boston, MA · Staff+

Sponsorship not specifiedDetected 166 days ago
Machine LearningElectrical Engineering

About the role

  • This role will be performed onsite from one of our offices in Santa Clara, CA or Boston, MA.
  • Evaluate end-to-end value proposition on ML workloads, translating model outputs into actionable guidance for the silicon team.
  • Work day-to-day with silicon architects, system designers, and workload owners to align performance expectations and constraints.

Responsibilities

  • You'll work from first principles, move quickly from insight to execution, and see your contributions directly reflected in what we build.
  • You'll work close with silicon architects and workload teams to explore design tradeoffs, validate performance assumptions, and identify bottlenecks early in the development cycle.
  • This role is well-suited for engineers who enjoy reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software evolve together.
  • Build and maintain system-level performance models for the rack-scale ML infrastructure with our custom silicon components.

Requirements

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or a closely related field.
  • 5-10+ years of experience in performance modeling for computer architectures, accelerators, or high-performance networking systems.
  • Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior.
  • Perfect understanding of ML systems: workload sharding, KV caching hierarchies, attention optimizations, trade-offs when deploying ML models at scale and various assumptions.
  • Ability to learn quick new ML architectures as soon as they come out, and build performance models for them.

Nice to have

  • PhD in Computer Science, Electrical Engineering, or a related field.
  • Prior experience modeling performance for ML accelerators and/or ML systems broadly.
  • Familiarity with shared memory systems and frameworks (e.g. CUDA VMM).
  • Experience with scale-up and high-bandwidth interconnects (e.g. NVLink or similar technologies).
  • Well-equipped, sunny offices in Santa Clara, CA and Boston, MA
  • Relocation assistance and visa sponsorship
  • The work we do here compounds across state-of-the-art AI models, systems, and real-world applications.
  • Timing: Joining now means real ownership of the company and meaningful influence over product direction and execution.

Compensation

  • Competitive salary commensurate with experience including base salary, performance-based bonus, and early stage equity grant
  • Well-equipped, sunny offices in Santa Clara, CA and Boston, MA
  • Relocation assistance and visa sponsorship

Benefits

  • Comprehensive benefits including health, dental, vision, and life insurance
  • Perks include a daily lunch stipend, 401k match, and more
  • A collaborative, continuous-learning work environment with smart, dedicated colleagues engaged in developing the next generation of architecture for high-performance computing

Company info

  • We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency.
  • You'll work alongside a group of people who care deeply about rigor, clarity, and impact.
  • We value thoughtful disagreement, fast learning, and intellectual fearlessness.
  • This is a place where strong ideas shine, curiosity is encouraged, and growth is a daily practice.
  • Impact: We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency.
  • We are seeking a Member of Technical Staff, Performance Modeling to develop performance models for end-to-end ML systems based on our silicon products.

Visa & Work Authorization

  • Relocation assistance and visa sponsorship

This listing is sourced directly from Netpreme's careers page and normalized into a canonical job model.