TBC

TBC

Research Scientist, Performance Engineering

San Francisco

Sponsorship not specifiedDetected 9 days ago
Machine LearningPyTorchLLMsRoboticsResearch

About the role

  • This role is focused on making frontier models run faster, cheaper, and more reliably - especially LLMs, diffusion models, video generation models, and world-model systems.
  • You will work across inference optimization, training efficiency, model compression, memory management, and GPU-level performance to help turn research systems into scalable, customer-ready products.

Responsibilities

  • Optimize inference for LLMs, diffusion models, video models, and world-model systems
  • Build and optimize high-throughput inference pipelines for large models running on GPU clusters
  • Implement custom kernels or low-level optimizations using Triton, CUDA, PyTorch, or related systems

Requirements

  • Hands-on experience optimizing LLMs, diffusion models, video generation models, or other large generative systems
  • Experience with one or more of:
  • Ability to reason about trade-offs between quality, latency, throughput, memory, and cost

Nice to have

  • Prior work optimizing large-scale generative models in production or research settings
  • Experience with modern inference/training stacks such as PyTorch, Triton, CUDA, vLLM, TensorRT, DeepSpeed, FSDP, Ray, or similar tooling
  • Experience working with LLMs, diffusion models, video generation models, or world models

This listing is sourced directly from TBC's careers page and normalized into a canonical job model.