Sciforium

Sciforium

GPU Kernel Engineer

San Francisco

Sponsorship not specifiedDetected 16 days ago
PythonC++Distributed SystemsFull-Stack DevelopmentMachine LearningPyTorchLLMsLLMOpsMLOpsSystems EngineeringElectrical EngineeringResearch

About the role

  • We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators.
  • You will work across the hardware-software stack, from low-level kernel development to integrating optimized ops into high-level ML frameworks used for large-scale training and inference.
  • This role is ideal for someone who thrives at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads, and who wants to make meaningful contributions to the efficiency and scalability of our ML platform.

Responsibilities

  • Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.
  • Profile and optimize end-to-end performance of ML operations, with a focus on large-scale LLM training and inference.
  • Develop performance models, identify bottlenecks, and deliver kernel-level improvements that significantly accelerate AI workloads.
  • Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack.

Requirements

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
  • Hands-on experience with Triton and/or JAX Pallas for custom kernel development.
  • Strong understanding of PTX, GPU ASM, and low-level GPU execution.
  • Experience with AMD GPUs and ROCm optimization.
  • Experience with efficient model serving frameworks (e.g., vLLM, TensorRT).
  • Experience with TPUs, XLA, or similar accelerator programming environments.
  • Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks.
  • Familiarity with JAX FFI and custom ML operator development.

Skills

  • Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes.
  • Contribute to tooling, documentation, benchmarking suites, and testing frameworks to ensure correctness and performance reproducibility.

Compensation

  • Competitive salary and equity

Benefits

  • Medical, dental, and vision insurance
  • Flexible time off
  • Competitive salary and equity

Equal opportunity

  • Sciforium is an equal opportunity employer.
  • All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Visa & Work Authorization

  • Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.