Sciforium

Sciforium

Senior HPC & GPU Infrastructure Engineer

San Francisco · Senior

Sponsorship not specifiedDetected 76 days ago
PythonBashNode.jsDistributed SystemsFull-Stack DevelopmentKubernetesTerraformAnsibleLinuxSite Reliability EngineeringMachine LearningPyTorchCybersecurityNetwork SecuritySystems EngineeringElectrical EngineeringFirewallResearch

About the role

  • We are seeking a Senior HPC & GPU Infrastructure Engineer to take full ownership of the health, reliability, and performance of our GPU compute cluster.
  • You will be the primary PyTOrchcustodian of our high-density accelerator environment and the linchpin between hardware operations, distributed systems, and machine learning workflows.
  • This role spans everything from hands-on Linux systems engineering and GPU driver bring-up to maintaining the ML software stack (CUDA/ROCm, PyTorch, JAX, vLLM).

Responsibilities

  • OS Management: Install, patch, and maintain Linux distributions (Ubuntu / CentOS / RHEL). Ensure consistent configuration, kernel tuning, and automation for large node fleets.
  • Identity & Storage Management: Manage LDAP/FreeIPA/AD for user identity, and administer distributed file systems such as NFS, GPFS, or Lustre.
  • Deployment & Bring-Up: Lead deployment of new GPU nodes, including BIOS configuration, NUMA tuning, GPU topology validation, and cluster integration.
  • Driver & Kernel Management: Build and optimize kernel modules, maintain GPU drivers and runtime stacks for both NVIDIA (CUDA) and AMD (ROCm).
  • Software Stack Maintenance: Maintain and optimize ML frameworks and libraries PyTorch, JAX, CUDA toolkit, cuDNN, ROCm, NCCL, and supporting runtime systems.

Requirements

  • 5+ years of experience in HPC, GPU cluster operations, Linux systems engineering, or similar roles.
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • Strong expertise with NVIDIA (H100/B200) or AMD (MI325x/MI355x) GPUs, including driver and kernel-level debugging.
  • Experience with network security, including VPNs, iptables/firewalld, SSH, and identity management (LDAP/FreeIPA/AD).
  • Proficiency in Bash and Python for scripting, automation, and workflow tooling.
  • Ideal candidate profile
  • Deep debugging experience with NVLink/NVSwitch fabrics and RDMA networking.

Nice to have

  • Experience with job schedulers such as Slurm, Kubernetes, or Run:AI.
  • Exposure to vLLM, model serving optimizations, or inference systems.
  • Hands-on experience with configuration management tools (Ansible, SaltStack, Terraform).
  • Previous experience supporting ML research teams in a startup or research-heavy environment.
  • Daily lunch, snacks, and beverages

Compensation

  • Competitive salary and equity

Benefits

  • Medical, dental, and vision insurance
  • Flexible time off
  • Competitive salary and equity
  • Implement and maintain monitoring for GPU health, thermal behavior, PCIe/NVLink topology issues, memory errors, and overall system load.
  • System Health & Reliability (SRE)

Equal opportunity

  • Sciforium is an equal opportunity employer.
  • All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Visa & Work Authorization

  • Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.