Perplexity AI
Member of Technical Staff (AI Inference Engineer)
San Francisco · Staff+
Sponsorship not specifiedDetected 99 days ago
PythonRustSassDistributed SystemsKubernetesMachine LearningDeep LearningTensorFlowPyTorchNLPLLMsResearchCommunication
About the role
- Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.
- Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. - Rust-native serving runtime.
- Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. - Reliability and observability.
Responsibilities
- We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.
Requirements
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
- Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
- Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
- Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).
Skills
- Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
- PyTorch internals, torch.compile, custom operators.
Company info
- New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
- GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
- Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
- Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
- Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
- Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
- You understand modern LLM architectures and are able to bring them up reliably in a production environment.
- You've built and operated production distributed systems under real load - ideally performance-critical ones.
Apply directly at Perplexity AI →Create a free account for alerts like thisView Perplexity AI immigration profile
This listing is sourced directly from Perplexity AI's careers page and normalized into a canonical job model.