Cerebras Systems

Cerebras Systems

Staff Software Engineer, GPU Inference

Toronto, CAN · Staff+

Sponsorship not specifiedDetected 39 days ago
PythonC++Node.jsDistributed SystemsKubernetesCI/CDLinuxMachine LearningPyTorchLLMsIncident ResponseComplianceElectrical EngineeringLeadershipCommunication

Stay score

odds of building a lasting career here

40Risky
Cap-exempt (no lottery)0
Sponsors this role90
Entry-level history0
PERM / green-card track0
Lottery odds40
Fits your clock70

Thin sponsorship signal and lottery-bound. A low-probability bet with your clock running. Prioritize cap-exempt roles and proven entry-level sponsors first.

Lottery odds assume a STEM candidate.

Personalize to your clock →

Employer immigration record

from this employer's Department of Labor filings

Green-card filing pattern in this occupation

Context, not a finding about this posting: of this employer's 3 green-card filings in this occupation, 100% were for a worker who already held the job.

Files H-1B transfers

22 transfer filings in the last year, covering 22 workers. Median labor-condition decision: 7 days. An employer that already files transfers is one that can take over an existing H-1B.

Sourced from Department of Labor LCA, PERM and prevailing-wage disclosure data. Employer matching is by name, so figures may be split across an employer's legal entities. Absence of a filing means none appears in our copy of the data, not that none exists.

Community outcomes

No reports yet — be the first to help the next applicant.

About the role

  • This is a hands-on role requiring deep debugging and optimization across application, runtime, distributed systems, and hardware layers.

Responsibilities

  • Productionize the GPU inference stack. Design, build, deploy, and maintain the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Drive reliability in production. Define service-level indicators and objectives for GPU-backed inference. Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.
  • Improve inference performance. Profile and optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under representative production workloads.
  • Optimize model-serving behavior. Tune and improve scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
  • Ensure numerical correctness. Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and compatibility across software and hardware releases.
  • Build performance and correctness infrastructure. Develop representative benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates. Turn one-off investigations into repeatable engineering systems.
  • Design, build, deploy, and maintain the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • You will write production code, establish operational practices for a new accelerator fleet, and drive improvements in time to first token, throughput, tail latency, and capacity efficiency.
  • Own GPU operational readiness.
  • Build automation that makes driver, firmware, runtime, model, and container compatibility explicit and reproducible.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Strong programming ability in C++ and Python, including experience with multithreading, concurrency, memory management, and performance-sensitive software.
  • Hands-on experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent internally developed system.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling methodology.
  • Experience debugging distributed systems across multiple layers rather than treating the serving framework or accelerator runtime as a black box.
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and operating latency-sensitive services in production.
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Nice to have

  • Experience with AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, Composable Kernel, or related libraries and tools.
  • Deep CUDA experience that demonstrates an ability to transfer GPU systems knowledge across accelerator platforms.
  • Experience modifying or contributing to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
  • Experience optimizing prefill-heavy or disaggregated prefill/decode inference architectures.
  • Understanding of KV-cache transfer, prefix caching, continuous batching, chunked prefill, request scheduling, and memory-aware admission control.
  • Experience with multi-GPU and multi-node inference, including tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
  • Experience optimizing Mixture-of-Experts or multimodal models.
  • Knowledge of GPU kernel optimization, operator fusion, graph capture, attention kernels, GEMM tuning, and communication/computation overlap.

Skills

  • Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.

Benefits

  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet.

Company info

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

This listing is sourced directly from Cerebras Systems's careers page and normalized into a canonical job model.