Mirendil

Mirendil

Member of Technical Staff, Inference

San Francisco · Staff+

Sponsorship not specified$300k-$500kDetected 28 days ago
LLMsResearch

About the role

  • You'll work across the full inference stack, from serving infrastructure down to hardware-level optimization.
  • Some example areas you might work on (not limited to):
  • If you're excited about pushing the performance limits of frontier model inference, we'd love to hear from you.

Responsibilities

  • Design and build high-throughput, low-latency inference serving systems for frontier models, optimizing for both research iteration and production deployment
  • Implement and validate inference-time optimizations: speculative decoding, quantization, KV cache management, and batching strategies
  • Build observability and reliability infrastructure so the team can measure latency, throughput, and cost across every serving configuration
  • Partner directly with teams to bring new model architectures and post-training techniques into production quickly

Compensation

  • We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Benefits

  • We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
  • Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.

Company info

  • We are building a frontier AI research company and training our own models end-to-end.
  • We are looking for an engineer to own the inference systems that power our models in production and research.

This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.