Biohub

Biohub

Staff HPC Engineer

San Francisco, CA (Hybrid) · Staff+

Sponsorship not specified$214k-$268kDetected 50 days ago
DockerKubernetesTerraformAnsibleMachine LearningTensorFlowPyTorchSystems EngineeringResearchExperimental Design

About the role

  • This hybrid environment will power everything from traditional HPC workloads to large AI training jobs, generative model development, real-time inference, and data-intensive pipelines.
  • The successful candidate will be a thought leader in HPC infrastructure, capable of partnering with scientists, computational biologists, and software engineers to translate complex research needs into high-impact computing solutions.
  • You will also foster adoption of emerging AI tools, and ensure our systems can scale to meet the demands of next-generation biomedical research.

Responsibilities

  • Architect and optimize environments for large-scale AI training and tuning, and low-latency scientific workloads.
  • Implement advanced resource scheduling and orchestration (Slurm, Kubernetes, SUNK) optimized for mixed HPC and AI workflows.
  • Support researchers with job optimization, GPU utilization best practices, and performance tuning for AI and HPC applications.
  • Evaluate, deploy, and maintain AI/ML software stacks (e.g., PyTorch, TensorFlow, Hugging Face, RAPIDS) and HPC toolchains.
  • Work with diverse science teams to translate research requirements into hardware/software solutions, from experimental design through publication.
  • 10+ years building and managing HPC infrastructure, with significant experience integrating AI/ML workloads.
  • Exceptional ability to collaborate with multidisciplinary teams and communicate complex technical concepts clearly.
  • Provides a generous employer match on employee 401(k) contributions to support planning for the future.
  • Relocation support for employees who need assistance moving

Requirements

  • Bachelor's or advanced degree in Computer Science, AI/ML, Data Science, Systems Engineering, or related field.
  • Strong hands-on experience with AI frameworks (PyTorch, TensorFlow, JAX) and distributed training strategies (Horovod, DeepSpeed, Ray).
  • Proficiency in automation tools (Terraform, Ansible, Puppet), containerization (Docker, Singularity), and orchestration frameworks.

Skills

  • Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere.
  • protein language models, genomic foundation models, and scientific reasoning systems built to be shared.
  • Our infrastructure supports day-to-day AI researcher workflows.
  • The Opportunity

Compensation

  • The San Francisco, CA base pay range for a new hire in this role is for a Staff HPC Engineer 214,000-$268,000 and for a Senior Staff HPC Engineer $241,000-$300,000.
  • New hires are typically hired into the lower portion of the range, enabling employee growth in the range over time.
  • Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process.
  • This position may be eligible to participate in our discretionary annual performance bonus program. Bonus eligibility and targets are determined in accordance with our total rewards philosophy and may vary by role.
  • Better Together
  • As we grow, we're excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment.

Benefits

  • Paid time off to volunteer at an organization of your choice.
  • Funding for select family-forming benefits.
  • Benefits for the Whole You
  • To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible.
  • Foster a culture of shared learning by running internal workshops on HPC-AI tooling (e.g., VS Code remote dev, containerization, MLOps workflows).

Company info

  • The HPC Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI.
  • We own the design, operation, and reliability of hybrid GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared.
  • The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization.

This listing is sourced directly from Biohub's careers page and normalized into a canonical job model.