Oscar Technology

Oscar Technology

Machine Learning Inference Engineer

San Francisco, California, USA · full-time

Sponsorship not specified$200k-$210kDetected 54 days ago
PythonKubernetesMachine LearningPyTorchNLPComputer VisionLLMsA/B TestingResearch

About the role

  • This is a highly technical and hands-on engineering role focused on production inference optimization for multimodal and generative AI systems.
  • The ideal candidate will have deep expertise in GPU inference, model serving, PyTorch-based deployment, and performance optimization for large-scale AI applications.

Responsibilities

  • model quality latency infrastructure cost production reliability Own the full inference optimization lifecycle from experimentation to production deployment Ideal Background Strong experience building and optimizing AI inference systems in production Deep understanding of
  • The role offers significant ownership across infrastructure, inference systems, and production model optimization, with opportunities to contribute to novel AI system design and scalable deployment architectures.

Compensation

  • $200k-$210k

Benefits

  • We are partnering with a fast-growing AI startup building next-generation multimodal generative systems focused on highly realistic visual experiences at scale.
  • The company operates at the intersection of computer vision, generative AI, and real-time inference infrastructure, developing advanced AI products used by enterprise customers across large consumer-facing industries.
  • diffusion models vision-language models (VLMs) large vision pipelines
  • Python PyTorch CUDA TensorRT Triton vLLM Experience with multimodal AI, computer vision, or generative AI systems Familiarity with diffusion models or large-scale vision pipelines is strongly preferred

This listing is sourced directly from Oscar Technology's careers page and normalized into a canonical job model.