Primeintellect
Member of Technical Staff - Training Platform
San Francisco · Staff+
Sponsorship not specified$150k-$300kDetected 14 days ago
TypeScriptPythonReactNext.jsTailwind CSSNode.jsFastAPIBackend DevelopmentFull-Stack DevelopmentAlgorithmsDatabricksGCPCloud PlatformsKubernetesTerraformAnsibleHelmLinuxPrometheusGrafanaDatadogAPI DevelopmentRESTLLMs
About the role
- The next generation of AI companies, enterprises, and research teams do not just need more GPUs.
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets
- Build and maintain Helm charts that compose trainers, inference servers, environment servers, and supporting services into reproducible "Training stacks"
- Develop the Python control-plane agents that watch pods, report run state to the platform, and keep clusters in sync
- Implement scheduling and autoscaling for heterogeneous hardware (H100/H200/B200) using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling
- Build node-local model caches, checkpoint pipelines, and shared storage for fast cold starts
- the infrastructure frontier AI labs build internally, made available to every ambitious AI team.
- You'll help build our hosted training platform - the product that lets users launch LoRA and full fine-tuning runs on managed GPU clusters with a single API call or a few clicks.
- Full visa sponsorship and relocation support
- Comfortable building Python control-plane agents that talk to Kubernetes APIs
Requirements
- Strong working knowledge of the modern AI stack - open model families, finetuning techniques (LoRA, QLoRA, full FT, RLHF/RLAIF), inference engines (vLLM, SGLang, TensorRT-LLM)
- Familiarity with GPU hardware tradeoffs (H100 / H200 / B200, NVLink, interconnects, memory hierarchy) and what they mean for training and inference workloads
- Awareness of what's happening at the frontier - new models, training methods, infra patterns - and the ability to translate that into product decisions
Nice to have
- Cloud platform experience (GCP preferred
Skills
- Understanding of distributed training fundamentals (data/tensor/pipeline/expert parallelism, NCCL, multi-node scheduling)
- Strong Kubernetes operations experience - Helm, CRDs, operators, KEDA, gang scheduling, GPU operator
- Comfortable debugging real production clusters (kubectl, pod lifecycle, node issues, networking)
- Cloud platform experience (GCP preferred - GCS, GKE, Cloud Run, Cloud Tasks)
- Infrastructure automation (Helm, Terraform, Ansible) and a GitOps mindset
- Observability: Prometheus, Grafana, Loki, OpenTelemetry, DCGM
- Linux fundamentals: networking, namespaces, performance tuning
- Strong Python backend development (FastAPI, async, SQLAlchemy)
- Modern frontend development (TypeScript, React/Next.js, Tailwind, shadcn) - enough to ship product surfaces end-to-end
- REST and tRPC API design
- Experience building developer tools, dashboards, and live-monitoring UIs
- Ship product UI in Next.js / React / TypeScript with shadcn, Tailwind, tRPC, and TanStack Query
Compensation
- Cash compensation $150K–$300K with significant equity
Benefits
- Cash compensation $150K-$300K with significant equity
- Flexible work arrangement (remote or San Francisco office)
Company info
- Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale - from SFT and RL to tool use, agent workflows, and continuously improving production models.
- Apply now and join us in our mission to make powerful AI models accessible to everyone.
Visa & Work Authorization
- Full visa sponsorship and relocation support
Apply directly at Primeintellect →Create a free account for alerts like thisView Primeintellect immigration profile
This listing is sourced directly from Primeintellect's careers page and normalized into a canonical job model.