Stack AV
Senior Software Engineer, Machine Learning Inference Platform
Pittsburgh, PA or Remote · Senior · Contract
Sponsorship not specifiedDetected 6 days ago
GoRustC++Distributed SystemsRESTgRPCWebSocketsMachine LearningDeep LearningPyTorchAccountingRoboticsProblem Solving
About the role
- You will be the go-to engineer for one or more areas such as model onboarding, serving APIs, metering, observability, performance optimization, or tenant isolation.
Responsibilities
- Own technical design and delivery of subsystems in a high-throughput, low-latency inference platform capable of handling multi-tenant, enterprise-grade inference workloads.
- Develop robust API layers (gRPC, WebSockets, REST, etc.) and developer SDKs that abstract complex distributed inference orchestration into seamless, reliable token streams.
- Build and harden a multi-tenant control plane to enable accurate metering, rate limiting, quotas, tenant isolation and noisy-neighbor fairness across the platform.
- Optimize inference performance across the entire system stack, including the model engine layer.
- Build observability and SLOs to gain insights into system economics, cache-hit rates, GPU utilization and cost accounting per model and per tenant.
- Partner with product and infrastructure teams on model onboarding, capacity planning, external API contracts and customer adoption.
- Decompose ambiguous work, drive issues to closure, and raise the engineering bar through code quality, reviews, testing, and mentoring.
- In the Senior Engineer role, you will own meaningful subsystems of Stack AV's inference platform and drive them from design through production.
Requirements
- Hands-on experience with large-scale inference services on GPUs, including KV caches, prefill/decode stages and throughput/latency trade-offs.
- Direct experience with inference engines (TensorRT, vLLM, etc) or serving frameworks (Dynamo, Triton or equivalent).
- Education: Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- Strong Data & ML systems fundamentals: data-intensive distributed systems, concurrency, networking and performance profiling.
- Technical Skills:
- We are proud to be an equal opportunity workplace. We believe that diverse teams produce the best ideas and outcomes. We are committed to building a culture of inclusion, entrepreneurship, and innovation across gender, race, age, sexual orientation, religion, disability, and identity.
Skills
- Strong programming skills in C++, Go, Rust or Python.
- Familiarity with deep learning frameworks (PyTorch, etc.) as well as model parallelism.
- Familiarity with GPU computing primitives such as CUDA, NCCL, NVLink, and hardware-specific optimizations.
- Practical understanding of high-performance networking architectures, including InfiniBand, RoCE, and low-latency cluster communication.
- Problem-Solving: Strong analytical and problem-solving skills.
- Autonomous vehicles (AV) experience is a bonus.
- This position may also involve working with software and technologies subject to U.S. export control regulations.
Visa & Work Authorization
- As such, this position may be contingent upon Stack AV verifying a candidate's residence, U.S. person status, and/or citizenship status
Apply directly at Stack AV →Create a free account for alerts like thisView Stack AV immigration profile
This listing is sourced directly from Stack AV's careers page and normalized into a canonical job model.