Cerebras Systems

Cerebras Systems

AI Engineer, Model Quality and Performance

Headquarters/Sunnyvale Office

Sponsorship not specifiedDetected 68 days ago
GitDockerAgentic AICompliance

About the role

  • We want someone whose first instinct is "how do I get an AI agent to do this on a loop."
  • You'll sit between engineering, product, and customer-facing teams.

Responsibilities

  • Design eval suites with AI agents in the loop. For every model release, curate a thoughtful mix of advanced, basic, long-context, and customer-use-case-specific evals. Use Claude to generate, validate, and prune candidate test cases at speed.
  • Build product-quality tooling that synthesizes quality + performance data into a single, easy-to-use view.
  • Experience building AI agents.
  • Design eval suites with AI agents in the loop.
  • A taste for tooling design.

Requirements

  • Comfort with Docker, Git, and the standard automation stack
  • Experience designing evals for agentic / coding / long-context / multimodal use cases.
  • Familiarity with open-source eval frameworks (EvalScope, lm-eval-harness, etc.).
  • Experience building AI agents. You ship real systems with Claude (or equivalent) as a force multiplier. You've built things that would have been infeasible solo without AI agents in the loop.
  • Strong math/stats background..
  • A taste for tooling design. You've shipped something that a non-engineer used without complaining. Bonus if AI helped you ship it.
  • Performance-tuning experience on custom silicon, GPUs, or FPGAs.

Nice to have

  • You've shipped something that a non-engineer used without complaining.

Skills

  • Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.

Benefits

  • Bonus if AI helped you ship it.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
  • You will define what "good" looks like across the models we serve, building AI-driven systems to measure it at scale, and translating those signals into artifacts our customers and product team actually use.
  • Build custom evals for target customers by orchestrating AI agents to mine trajectories from their workloads and synthesize representative eval sets.
  • Build automations to forecast and benchmark model performance on Cerebras for our top customers, including modeling how fast customer-specific workloads will run in production.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.