Mirendil
Member of Technical Staff, Post-Training, RL Environments
San Francisco · Staff+
Sponsorship not specified$300k-$500kDetected 28 days ago
Design SystemsResearch
About the role
- Some example areas you might work on (not limited to):
Responsibilities
- The quality of our models depends directly on the quality of the data and environments we train on; you will own those systems end-to-end.
- Build and automate data collection pipelines for complex, long-horizon RL tasks.
- Build robust systems to identify and prevent reward hacking.
- Build scalable sandboxed execution environments for realistic tasks involving potentially multiple agents, nodes, and users.
- Design systems to estimate the influence of training environments on production model behavior.
- Collaborate with teams across the stack to identify potential axes of improvements in production model behavior, and develop training environments to push these axes.
Compensation
- We offer a base salary of $300,000–$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
Benefits
- We offer a base salary of $300,000-$500,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.
- Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments.
- We are looking for a research engineer to build the data systems and execution environments that power reinforcement learning at Mirendil.
Company info
- We are building a frontier AI research company and training our own models end-to-end.
Apply directly at Mirendil →Create a free account for alerts like thisView Mirendil immigration profile
This listing is sourced directly from Mirendil's careers page and normalized into a canonical job model.