Sciforium

Sciforium

Lead Software Engineer, Model Serving Platform

San Francisco

Sponsorship not specifiedDetected 228 days ago
PythonC++Full-Stack DevelopmentCode ReviewKubernetesMachine LearningLLMsMLOpsSystems EngineeringElectrical EngineeringResearchCommunication

About the role

  • You will learn and shape the full AI stack: from GPU kernels and quantized execution paths to distributed serving, scheduling, and the APIs that power real-time AI applications.

Responsibilities

  • Lead the technical direction of the model serving platform, owning architecture decisions and guiding engineering execution.
  • Build core serving components including execution runtimes, batching, scheduling, and distributed inference systems.
  • Develop high-performance C++ and CUDA/HIP modules, including custom GPU kernels and memory-optimized runtimes.
  • Collaborate with ML researchers to productionize new multimodal models and ensure low-latency, scalable inference.
  • Build Python APIs and services that expose model capabilities to downstream applications.
  • Mentor and support other engineers through code reviews, design discussions, and hands-on technical guidance.
  • Drive performance profiling, benchmarking, and observability across the inference stack.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
  • 5+ years of experience designing and building scalable, reliable backend systems or distributed infrastructure.
  • Strong understanding of LLM inference mechanics (prefill vs decode, batching, KV cache)
  • Experience with Kubernetes/Ray, Containerization
  • Strong proficiency in C++, Python.

Nice to have

  • Experience with ML systems engineering, distributed GPU scheduling, open source inference engine like vLLM, Sglang, or TRT-LLM
  • Proficiency in CUDA or ROCm and experience with GPU profiling tools
  • Experience at an AI/ML startup, research lab, or Big Tech infrastructure/ML team.
  • Familiarity with multimodal model architectures, raw-byte models, or efficient inference techniques.
  • Contributions to open-source ML or HPC infrastructure
  • Daily lunch, snacks, and beverages

Compensation

  • Competitive salary and equity

Benefits

  • Medical, dental, and vision insurance
  • Flexible time off
  • Competitive salary and equity

Equal opportunity

  • Sciforium is an equal opportunity employer.
  • All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Visa & Work Authorization

  • Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

This listing is sourced directly from Sciforium's careers page and normalized into a canonical job model.