Bespokelabs

Bespokelabs

RL Environments

Mountain View

Sponsorship not specifiedDetected 337 days ago
PythonAWSGCPCloud PlatformsMachine LearningPyTorchData AnalysisLLMsResearch

About the role

  • This role combines research intuition with practical execution.
  • You'll need to understand agent behavior deeply-spotting reward hacking, analyzing failure modes, and diagnosing why certain environments produce better training outcomes.
  • Then you'll translate that understanding into repeatable processes and benchmark suites that we can showcase externally.

Responsibilities

  • Develop systematic strategies and recipes for creating high-quality RL environments that effectively train and evaluate agents.
  • Study how LLMs and agents fail across different task types, identifying patterns that inform better environment design.
  • Create benchmark environments that test specific agent capabilities, packaging them for external release on our evaluation platform.
  • Collaborate with the team to ensure benchmarks integrate smoothly into our external-facing dashboards.
  • Establish quality standards and evaluation protocols that maintain high bars as we scale environment production.
  • You'll develop systematic approaches to environment design, identify where agents fail, and turn those insights into high-quality training data and benchmarks.
  • If you're passionate about understanding agent behavior and creating systematic approaches to environment design, we'd love to hear from you.

Nice to have

  • Experience with AI safety, robustness testing, or adversarial evaluation
  • Publications or projects related to RL, agent evaluation, or data-centric AI
  • Experience shipping research artifacts (datasets, benchmarks, evaluation suites) to the community

Skills

  • Strong foundation in machine learning-either through a PhD/MS in ML, CS, or equivalent industry experience.
  • Deep curiosity about agent behavior and failure modes, with ability to form hypotheses and test them systematically.
  • Experience analyzing complex systems and extracting actionable insights from data.
  • Patience and attention to detail for studying agent rollouts and identifying subtle patterns.
  • Technical execution:
  • Proficiency in Python and ML frameworks (PyTorch, JAX, or similar).
  • Experience with RL concepts and agent training, even if not from a RL background.
  • Comfortable working with cloud platforms (GCP, AWS) for running experiments at scale.
  • Practical engineering:
  • Experience with data analysis tools and creating reproducible workflows.
  • Systematic approach to quality verification and testing.
  • Hands-on experience with reinforcement learning or agent training systems

Compensation

  • Competitive salary and equity based on experience and background
  • We encourage applications from candidates with diverse research backgrounds. If you're passionate about understanding agent behavior and creating systematic approaches to environment design, we'd love to hear from you.

Benefits

  • Health coverage, flexible work arrangements, and the opportunity to shape how the AI community evaluates and trains agents

This listing is sourced directly from Bespokelabs's careers page and normalized into a canonical job model.