SpreeAI

SpreeAI

Principal Engineer, AI Platform & Infrastructure

Hybrid (San Francisco, California, US) · Principal

Sponsorship not specifiedDetected 20 days ago
PythonDistributed SystemsAWSGCPAzureCloud PlatformsDockerKubernetesCI/CDDevOpsPlatform EngineeringMachine LearningPyTorchAirflowLLMsMLOpsAI OrchestrationResearch

About the role

  • This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability.
  • Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems.
  • Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation.

Responsibilities

  • You will own critical infrastructure decisions across deployment, observability, and resource management, with direct impact on production systems serving real partner traffic.
  • We bring together cutting-edge AI and real-world retail to deliver production systems that redefine how people shop online.
  • What You'll Own ML Platform & Training Enablement Build and operate SPREEAI's end-to-end ML platform spanning training, evaluation, deployment, and monitoring.
  • Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability.
  • Standardize and support serving infrastructure using modern inference runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, Ray Serve, TorchServe, ONNX Runtime, or equivalent systems.
  • GPU Infrastructure & System Efficiency Design and manage GPU allocation, scheduling, and resource utilization across training and inference workloads.
  • Design and operate model evaluation and benchmarking systems, including automated regression detection and quality gates for production releases.
  • Partner with research teams to productionize new capabilities by providing robust infrastructure, tooling, and deployment pathways.
  • We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact.
  • About the Role SPREEAI is building the future of AI-powered commerce through photorealistic virtual try-on and multimodal intelligence.

Requirements

  • 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
  • Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads.
  • Strong understanding of distributed systems and large-scale ML infrastructure design.
  • Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow.
  • Experience deploying and managing production inference systems using platforms like Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services.
  • Strong understanding of inference optimization techniques such as batching, quantization, CUDA graphs, and memory-aware scheduling.
  • Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring.
  • Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers.

Nice to have

  • Experience with large-scale GPU clusters e.g. A100/H100, NCCL, and high-throughput data pipelines.
  • Experience designing evaluation and monitoring systems for generative AI workloads.
  • Familiarity with ML security, privacy, and data governance practices.
  • Why This Role Matters This is not a traditional DevOps role.
  • You will define the systems powering real-time AI experiences where latency, cost, and model quality directly impact end-user experience.

Skills

  • Reduce manual model deployment friction through standardized pipelines and tooling.
  • Improve GPU utilization and reduce training and inference costs.
  • Establish robust observability and evaluation gates for production model releases.
  • This is the infrastructure backbone that enables SPREEAI to turn frontier AI research into reliable, scalable, production-grade systems.
  • Why Join SPREEAI?

Company info

  • Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology.
  • What We're Looking For 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.

This listing is sourced directly from SpreeAI's careers page and normalized into a canonical job model.