SpreeAI
Principal Engineer, AI Platform & Infrastructure
Hybrid (San Francisco, California, US) · Principal
Sponsorship not specifiedDetected 20 days ago
PythonDistributed SystemsAWSGCPAzureCloud PlatformsDockerKubernetesCI/CDDevOpsPlatform EngineeringMachine LearningPyTorchAirflowLLMsMLOpsAI OrchestrationResearch
About the role
- This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability.
- Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems.
- Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation.
Responsibilities
- You will own critical infrastructure decisions across deployment, observability, and resource management, with direct impact on production systems serving real partner traffic.
- We bring together cutting-edge AI and real-world retail to deliver production systems that redefine how people shop online.
- What You'll Own ML Platform & Training Enablement Build and operate SPREEAI's end-to-end ML platform spanning training, evaluation, deployment, and monitoring.
- Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability.
- Standardize and support serving infrastructure using modern inference runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, Ray Serve, TorchServe, ONNX Runtime, or equivalent systems.
- GPU Infrastructure & System Efficiency Design and manage GPU allocation, scheduling, and resource utilization across training and inference workloads.
- Design and operate model evaluation and benchmarking systems, including automated regression detection and quality gates for production releases.
- Partner with research teams to productionize new capabilities by providing robust infrastructure, tooling, and deployment pathways.
- We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact.
- About the Role SPREEAI is building the future of AI-powered commerce through photorealistic virtual try-on and multimodal intelligence.
Requirements
- 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
- Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads.
- Strong understanding of distributed systems and large-scale ML infrastructure design.
- Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow.
- Experience deploying and managing production inference systems using platforms like Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services.
- Strong understanding of inference optimization techniques such as batching, quantization, CUDA graphs, and memory-aware scheduling.
- Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring.
- Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers.
Nice to have
- Experience with large-scale GPU clusters e.g. A100/H100, NCCL, and high-throughput data pipelines.
- Experience designing evaluation and monitoring systems for generative AI workloads.
- Familiarity with ML security, privacy, and data governance practices.
- Why This Role Matters This is not a traditional DevOps role.
- You will define the systems powering real-time AI experiences where latency, cost, and model quality directly impact end-user experience.
Skills
- Reduce manual model deployment friction through standardized pipelines and tooling.
- Improve GPU utilization and reduce training and inference costs.
- Establish robust observability and evaluation gates for production model releases.
- This is the infrastructure backbone that enables SPREEAI to turn frontier AI research into reliable, scalable, production-grade systems.
- Why Join SPREEAI?
Company info
- Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology.
- What We're Looking For 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
Apply directly at SpreeAI →Create a free account for alerts like thisView SpreeAI immigration profile
This listing is sourced directly from SpreeAI's careers page and normalized into a canonical job model.