Inception
Member of Technical Staff, Inference & Serving
San Mateo, USA · Staff+
Sponsorship not specifiedDetected 134 days ago
Distributed SystemsAWSGCPAzureDockerKubernetesCI/CDMachine LearningTensorFlowPyTorchAirflowLLMsAI OrchestrationLoad Balancing
About the role
- The Role We're looking for engineers and scientists to design, optimize, and scale the systems that power our diffusion LLMs in production. Your work will make inference faster, more cost-effective, and more reliable. Key Responsibilities Build and optimize high-performance model serving systems for low-latency inference of diffusion LLMs. Extend
- orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving. Implement and manage load balancing, autoscaling, and traffic routing for model endpoints. Build systems for model versioning, canary deployments, and zero-downtime rollouts. Develop monitoring, alerting, and observability tooling to ensure SLA
Responsibilities
- Build and optimize high-performance model serving systems for low-latency inference of diffusion LLMs.
- Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
- Build systems for model versioning, canary deployments, and zero-downtime rollouts.
- Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.
- Collaborate with ML researchers to translate model advances (new architectures, quantization techniques, batching strategies) into production-ready serving improvements.
Requirements
- Experience with model optimization techniques (quantization, distillation, speculative decoding, continuous batching).
Nice to have
- Your work will make inference faster, more cost-effective, and more reliable.
- Extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving.
- BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience).
- Knowledge of ML serving frameworks (SGLang, vLLM, Triton Inference Server, TensorRT-LLM).
- Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective.
- Familiarity with high-performance computing and GPU programming (CUDA).
- Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines.
- Background in performance optimization and profiling of ML systems.
Apply directly at Inception →Create a free account for alerts like thisView Inception immigration profile
This listing is sourced directly from Inception's careers page and normalized into a canonical job model.