Krea AI

Krea AI

Engineer, Supercomputing & Distributed Systems

San Francisco

Sponsorship not specifiedDetected 110 days ago
PythonExpressDistributed SystemsSQLKubernetesKafkaMachine LearningPyTorchPandasNumPyData EngineeringLLMsResearch

About the role

  • We believe AI is a new medium that allows us to express ourselves through various formats-text, images, video, sound, and even 3D.
  • You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc.

Responsibilities

  • Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
  • Manage distributed training and inference on 1000+ GPU Kubernetes clusters
  • Profile and optimize dataloaders streaming thousands of images per second
  • Build fault tolerance systems for large-scale pretraining
  • Collaborate with researchers on evolving RL infrastructure
  • Build the systems that bridge raw cluster capacity and research output

Requirements

  • Fundamental knowledge of containerization, operating systems, file-systems, and networking

Skills

  • Supercomputing / AI Infra at Krea
  • Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc.

This listing is sourced directly from Krea AI's careers page and normalized into a canonical job model.