Wizardcommerce

Wizardcommerce

Senior Machine Learning Engineer (Inference Platform)

Remote - USA · Senior

Sponsorship not specifiedDetected 48 days ago
PythonAWSGCPAzureCloud PlatformsCI/CDDevOpsPlatform EngineeringMachine LearningData ScienceLLMsMLOpsA/B TestingLeadership

About the role

  • This is not a traditional MLOps role focused solely on pipelines and tooling.
  • You'll be responsible for the inference infrastructure powering a live conversational shopping agent, operating multiple specialized serving engines under real-world production load.

Responsibilities

  • Own and evolve our multi-engine inference platform, supporting a variety of model types and serving requirements.
  • Build and improve production ML pipelines - taking models from experimentation to reliable, high-throughput serving.
  • Define and implement model versioning, rollout, rollback, and lifecycle management strategies that ensure reproducibility and operational reliability.
  • Build observability, monitoring, alerting, and operational tooling for production inference systems.
  • Optimize inference performance through efficient resource utilization, hardware-aware serving strategies, and cost-conscious infrastructure design.
  • Partner with ML, Data, Product, and DevOps teams to turn ideas into production systems, driving the technical decisions on serving and scale.

Requirements

  • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience.
  • 5-8+ years of experience in Software Engineering, ML Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of production ML serving systems.
  • Hands-on experience running an LLM serving engine (vLLM, TGI, TensorRT-LLM, or SGLang) in production under real load - not just managed or hosted endpoints.
  • Experience with cloud platforms such as AWS, GCP, or Azure, and familiarity with ML lifecycle tooling, experimentation platforms, and model registries.
  • Experience serving heterogeneous workloads, including LLMs, embedding models, and extraction models, each with distinct latency, throughput, and scaling requirements.
  • Demonstrated ability to balance latency, throughput, reliability, and infrastructure cost while operating production-scale ML systems.
  • Experience in high-growth startup environments and comfort operating in fast-moving, evolving technical landscapes.

Skills

  • About Wizard AI

Company info

  • What We're Looking For

This listing is sourced directly from Wizardcommerce's careers page and normalized into a canonical job model.