Mirendil
Member of Technical Staff, Inference
San Francisco · Staff+
Sponsorship not specified$300k-$500kDetected 28 days ago
LLMsResearch
About the role
- You'll work across the full inference stack, from serving infrastructure down to hardware-level optimization.
- Some example areas you might work on (not limited to):
- If you're excited about pushing the performance limits of frontier model inference, we'd love to hear from you.
Responsibilities
- Design and build high-throughput, low-latency inference serving systems for frontier models, optimizing for both research iteration and production deployment
- Implement and validate inference-time optimizations: speculative decoding, quantization, KV cache management, and batching strategies
- Build observability and reliability infrastructure so the team can measure latency, throughput, and cost across every serving configuration
- Partner directly with teams to bring new model architectures and post-training techniques into production quickly
Compensation
- We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Benefits
- We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
- Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
Company info
- We are building a frontier AI research company and training our own models end-to-end.
- We are looking for an engineer to own the inference systems that power our models in production and research.
Apply directly at Mirendil →Create a free account for alerts like thisView Mirendil immigration profile
This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.