Cerebras Systems
AI Inference Core - SDET Technical Lead, Release Integration Testing
Sunnyvale, CA
Stay score
odds of building a lasting career here
Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.
Lottery odds assume a STEM candidate.
Personalize to your clock →Employer immigration record
from this employer's Department of Labor filings
Green-card filing pattern in this occupation
Files H-1B transfers
Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.
Community outcomes
No reports yet — be the first to help the next applicant.
About the role
- The Production Engine for Inference Core - turning integrated features into reliable production releases.
- You will define the quality strategy across the pre-release and release cycle, from feature and model integration through branch stability, release qualification, deployment, and post-release learning.
- You will work across AI frameworks, runtime, compiler, kernels, distributed systems, infrastructure, and hardware to make release risk visible and actionable.
Responsibilities
- Own the inference-path readiness gate by reviewing unit, simulation, benchmark, feature-test, and integration evidence, documenting gaps, and approving integration readiness before release entry.
- Lead integrated inference E2E validation across features and the cloud-to-wafer stack
- Partner with and mentor SDETs, feature teams, Integration, Core Infra, release owners, and deployment teams
Requirements
- Experience designing automation and test architecture for distributed, systems-level, infrastructure, or AI software.
- Strong software-engineering fundamentals and programming ability in Python Go, or a similar language.
- Demonstrated technical leadership in software quality, test infrastructure, systems validation, release engineering, or complex software integration.
- Proven ability to break down ambiguous cross-stack failures, form hypotheses, gather evidence, and drive issues to resolution.
- Strong understanding of risk-based testing, release readiness, regression strategy, failure analysis, and quality metrics.
- Ability to influence and align multiple engineering teams without relying solely on organizational authority.
- Clear communication and sound judgment during high-pressure release situations, including the ability to explain technical risk to engineering and leadership audiences.
Skills
- Experience with AI infrastructure, model deployment, LLMs, multimodal workloads, or large-scale compute clusters.
- Experience with performance testing, profiling, observability, fault injection, reliability, or production failure analysis.
- Experience in a startup or similarly fast-moving, resource-constrained engineering environment.
- Track record of taking a quality or release capability from zero to one and scaling it across teams.
- Familiarity with containers, cluster orchestration, cloud infrastructure, CI/CD, or high-performance computing.
- Release readiness is based on explicit criteria and high-signal evidence rather than intuition.
- Fewer inference-path integration defects are first discovered in final release qualification or production.
- Cross-component risks are found earlier, debug cycles are shorter, and coverage ownership is explicit.
- Master and release-branch health is measurable, actionable, and steadily improving.
- Test automation and release infrastructure shorten feedback loops without sacrificing signal quality.
Benefits
- Define the Release Integration Testing strategy, engagement criteria, ownership boundaries, entry and exit criteria, coverage expectations, and escalation thresholds for AI Inference Core.
- Improve master and release-branch stability through actionable health metrics, failure classification, release-quality reporting, dashboards, qualification workflows, and release pipelines.
- Lead first-pass regression and rollout triage, coordinate owners through resolution, drive RCA, place missing coverage at the correct layer, and plan rollout across multiple product and release projects.
Company info
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.