Mirendil
Member of Technical Staff, Post-Training, RL Infra
San Francisco · Staff+
Sponsorship not specified$300k-$500kDetected 28 days ago
AlgorithmsResearch
About the role
- This role sits at the intersection of research and infrastructure.
- You will work to push the scale of our RL stack, whether it is novel recipe ideas, reliability, or performance.
- Some example areas you might work on (not limited to):
Responsibilities
- Design and build reliable infrastructure for large-scale RL training
- Implement novel performance optimizations across the training stack
- Develop evaluation and benchmarking infrastructure to measure model progress, throughput, and uptime
- Build data collection and feedback pipelines that close the loop between human signal, reward modeling, and training
- Collaborate with multiple teams to rapidly iterate on RL algorithms and get experiments into production training runs
Compensation
- We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Benefits
- We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
- Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
Company info
- We are building a frontier AI research company and training our own models end-to-end.
- We are looking for engineers to help build the post-training stack for frontier reasoning models.
Apply directly at Mirendil →Create a free account for alerts like thisView Mirendil immigration profile
This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.