Primeintellect

Primeintellect

Member of Technical Staff - Training Platform

San Francisco · Staff+

Sponsorship not specified$150k-$300kDetected 14 days ago
TypeScriptPythonReactNext.jsTailwind CSSNode.jsFastAPIBackend DevelopmentFull-Stack DevelopmentAlgorithmsDatabricksGCPCloud PlatformsKubernetesTerraformAnsibleHelmLinuxPrometheusGrafanaDatadogAPI DevelopmentRESTLLMs

About the role

  • The next generation of AI companies, enterprises, and research teams do not just need more GPUs.

Responsibilities

  • Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets
  • Build and maintain Helm charts that compose trainers, inference servers, environment servers, and supporting services into reproducible "Training stacks"
  • Develop the Python control-plane agents that watch pods, report run state to the platform, and keep clusters in sync
  • Implement scheduling and autoscaling for heterogeneous hardware (H100/H200/B200) using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling
  • Build node-local model caches, checkpoint pipelines, and shared storage for fast cold starts
  • the infrastructure frontier AI labs build internally, made available to every ambitious AI team.
  • You'll help build our hosted training platform - the product that lets users launch LoRA and full fine-tuning runs on managed GPU clusters with a single API call or a few clicks.
  • Full visa sponsorship and relocation support
  • Comfortable building Python control-plane agents that talk to Kubernetes APIs

Requirements

  • Strong working knowledge of the modern AI stack - open model families, finetuning techniques (LoRA, QLoRA, full FT, RLHF/RLAIF), inference engines (vLLM, SGLang, TensorRT-LLM)
  • Familiarity with GPU hardware tradeoffs (H100 / H200 / B200, NVLink, interconnects, memory hierarchy) and what they mean for training and inference workloads
  • Awareness of what's happening at the frontier - new models, training methods, infra patterns - and the ability to translate that into product decisions

Nice to have

  • Cloud platform experience (GCP preferred

Skills

  • Understanding of distributed training fundamentals (data/tensor/pipeline/expert parallelism, NCCL, multi-node scheduling)
  • Strong Kubernetes operations experience - Helm, CRDs, operators, KEDA, gang scheduling, GPU operator
  • Comfortable debugging real production clusters (kubectl, pod lifecycle, node issues, networking)
  • Cloud platform experience (GCP preferred - GCS, GKE, Cloud Run, Cloud Tasks)
  • Infrastructure automation (Helm, Terraform, Ansible) and a GitOps mindset
  • Observability: Prometheus, Grafana, Loki, OpenTelemetry, DCGM
  • Linux fundamentals: networking, namespaces, performance tuning
  • Strong Python backend development (FastAPI, async, SQLAlchemy)
  • Modern frontend development (TypeScript, React/Next.js, Tailwind, shadcn) - enough to ship product surfaces end-to-end
  • REST and tRPC API design
  • Experience building developer tools, dashboards, and live-monitoring UIs
  • Ship product UI in Next.js / React / TypeScript with shadcn, Tailwind, tRPC, and TanStack Query

Compensation

  • Cash compensation $150K–$300K with significant equity

Benefits

  • Cash compensation $150K-$300K with significant equity
  • Flexible work arrangement (remote or San Francisco office)

Company info

  • Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models.
  • Apply now and join us in our mission to make powerful AI models accessible to everyone.

Visa & Work Authorization

  • Full visa sponsorship and relocation support

This listing is sourced directly from Primeintellect's careers page and normalized into a canonical job model.