Bespokelabs
RL Environments
Mountain View
Sponsorship not specifiedDetected 337 days ago
PythonAWSGCPCloud PlatformsMachine LearningPyTorchData AnalysisLLMsResearch
About the role
- This role combines research intuition with practical execution.
- You'll need to understand agent behavior deeply-spotting reward hacking, analyzing failure modes, and diagnosing why certain environments produce better training outcomes.
- Then you'll translate that understanding into repeatable processes and benchmark suites that we can showcase externally.
Responsibilities
- Develop systematic strategies and recipes for creating high-quality RL environments that effectively train and evaluate agents.
- Study how LLMs and agents fail across different task types, identifying patterns that inform better environment design.
- Create benchmark environments that test specific agent capabilities, packaging them for external release on our evaluation platform.
- Collaborate with the team to ensure benchmarks integrate smoothly into our external-facing dashboards.
- Establish quality standards and evaluation protocols that maintain high bars as we scale environment production.
- You'll develop systematic approaches to environment design, identify where agents fail, and turn those insights into high-quality training data and benchmarks.
- If you're passionate about understanding agent behavior and creating systematic approaches to environment design, we'd love to hear from you.
Nice to have
- Experience with AI safety, robustness testing, or adversarial evaluation
- Publications or projects related to RL, agent evaluation, or data-centric AI
- Experience shipping research artifacts (datasets, benchmarks, evaluation suites) to the community
Skills
- Strong foundation in machine learning-either through a PhD/MS in ML, CS, or equivalent industry experience.
- Deep curiosity about agent behavior and failure modes, with ability to form hypotheses and test them systematically.
- Experience analyzing complex systems and extracting actionable insights from data.
- Patience and attention to detail for studying agent rollouts and identifying subtle patterns.
- Technical execution:
- Proficiency in Python and ML frameworks (PyTorch, JAX, or similar).
- Experience with RL concepts and agent training, even if not from a RL background.
- Comfortable working with cloud platforms (GCP, AWS) for running experiments at scale.
- Practical engineering:
- Experience with data analysis tools and creating reproducible workflows.
- Systematic approach to quality verification and testing.
- Hands-on experience with reinforcement learning or agent training systems
Compensation
- Competitive salary and equity based on experience and background
- We encourage applications from candidates with diverse research backgrounds. If you're passionate about understanding agent behavior and creating systematic approaches to environment design, we'd love to hear from you.
Benefits
- Health coverage, flexible work arrangements, and the opportunity to shape how the AI community evaluates and trains agents
Apply directly at Bespokelabs →Create a free account for alerts like thisView Bespokelabs immigration profile
This listing is sourced directly from Bespokelabs's careers page and normalized into a canonical job model.