Krea AI
Engineer, Supercomputing & Distributed Systems
San Francisco
Sponsorship not specifiedDetected 110 days ago
PythonExpressDistributed SystemsSQLKubernetesKafkaMachine LearningPyTorchPandasNumPyData EngineeringLLMsResearch
About the role
- We believe AI is a new medium that allows us to express ourselves through various formats-text, images, video, sound, and even 3D.
- You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc.
Responsibilities
- Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
- Manage distributed training and inference on 1000+ GPU Kubernetes clusters
- Profile and optimize dataloaders streaming thousands of images per second
- Build fault tolerance systems for large-scale pretraining
- Collaborate with researchers on evolving RL infrastructure
- Build the systems that bridge raw cluster capacity and research output
Requirements
- Fundamental knowledge of containerization, operating systems, file-systems, and networking
Skills
- Supercomputing / AI Infra at Krea
- Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc.
Apply directly at Krea AI →Create a free account for alerts like thisView Krea AI immigration profile
This listing is sourced directly from Krea AI's careers page and normalized into a canonical job model.