Inception

Inception

Member of Technical Staff, Backend, LLM Applications

San Mateo, USA · Staff+

Sponsorship not specifiedDetected 134 days ago
PythonDistributed SystemsBackend DevelopmentAWSAzureCloud PlatformsKubernetesTerraformCI/CDPrometheusGrafanaMachine LearningLLMsLoad Balancing

About the role

  • The Role We seek experienced backend engineers to own the systems that serve our diffusion LLMs in production. You'll build and operate infrastructure that handles billions of inference requests — optimizing for latency, throughput, cost, and reliability. This role sits at the intersection of ML systems and backend infrastructure. Key Responsibilities
  • Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs. Implement and manage load balancing, autoscaling, and traffic routing for model endpoints. Build systems for model versioning, canary deployments, and zero-downtime rollouts. Develop monitoring, alerting, and observability tooling to ensure SLA

Responsibilities

  • Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs.
  • Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
  • Build systems for model versioning, canary deployments, and zero-downtime rollouts.
  • Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.

Requirements

  • Experience with monitoring and observability tools (Prometheus, Grafana).

Nice to have

  • This role sits at the intersection of ML systems and backend infrastructure.
  • Benchmark and evaluate serving frameworks and hardware configurations to inform infrastructure decisions.
  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • 5+ years of experience building production backend systems.
  • Strong proficiency in Python, including async programming and concurrent systems.
  • Solid understanding of distributed systems, networking, and load balancing at scale.
  • Familiarity with Kubernetes, CI/CD pipelines, and cloud infra (AWS and/or Azure).
  • Preferred Skills Experience serving LLMs or other large generative models in production at scale.

This listing is sourced directly from Inception's careers page and normalized into a canonical job model.