Perplexity

Perplexity

Member of Technical Staff (AI Inference Engineer)

San Francisco · Staff+

Sponsorship not specifiedDetected 99 days ago
PythonRustSassDistributed SystemsKubernetesMachine LearningDeep LearningTensorFlowPyTorchNLPLLMsResearchCommunication

About the role

  • Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer to join us.
  • Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow. - Rust-native serving runtime.
  • Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving. - Reliability and observability.

Responsibilities

  • We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.

Requirements

  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation).

Skills

  • Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • PyTorch internals, torch.compile, custom operators.

Company info

  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents.
  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.

This listing is sourced directly from Perplexity's careers page and normalized into a canonical job model.