Labelbox
Forward Deployed Engineer, RL Environments
San Francisco Bay Area · Mid
Sponsorship not specified$140k-$200kDetected 106 days ago
PythonGoRustC++BashAWSGCPDockerCI/CDMachine LearningAgentic AIControlsTest AutomationResearch
About the role
- You'll write production-quality infrastructure code, integrate with open-source RL tooling, and work closely with our data operations team to ensure environments are robust, observable, and ready for human annotators and model agents alike.
- You won't be doing ML research, but you'll need to deeply understand how RL training loops consume environments and where the bottlenecks live.
Responsibilities
- Design, build, and maintain sandboxed RL environments for agentic AI training-including terminal emulators, browser automation harnesses, computer-use simulators, and tool-augmented workspaces (e.g., environments built on frameworks like TerminalBench, OSWorld, and Tau-bench)
- Develop reproducible, containerized execution environments (Docker, VMs, lightweight sandboxes) that support deterministic task rollouts and reward signal collection
- Build instrumentation and observability layers-structured logging, trajectory capture, state snapshotting-so training runs and human annotation sessions produce clean, auditable data
- Collaborate with data operations to design task curricula and evaluation protocols that stress-test model capabilities across environment types
- Own environment deployment and reliability: CI/CD pipelines, automated testing of environment configurations, and monitoring for drift or breakage across versions
- Experience building or maintaining developer tooling, CLI tools, or infrastructure automation
- Innovation at Speed: We celebrate those who take ownership, move fast, and deliver impact.
- We empower people to drive results through clear ownership and metrics.
- Ability to read and implement from academic papers and open-source benchmark repositories without extensive hand-holding
- Familiarity with GCP or AWS infrastructure (Compute Engine, ECS/EKS, Cloud Build)
Requirements
- 2+ years of professional software engineering experience, with strong fundamentals in Python and at least one systems-level language (Go, Rust, C++)
- Demonstrated experience with containerization and sandboxing (Docker, Podman, Firecracker, or similar) in production or near-production contexts
- Familiarity with RL concepts: MDPs, reward shaping, episode structure, observation/action spaces.
- Comfort working with browser automation frameworks or terminal interaction tooling
- Strong debugging instincts-you can trace failures across process boundaries, container layers, and network calls
- Experience with agentic AI evaluation frameworks (SWE-bench, WebArena, OSWorld, TerminalBench, or similar)
- You move fast, but you care about reliability because you know environments that break silently poison training data.
Skills
- Advanced annotation tools, workflow automation, and quality control systems that enable teams to produce high-quality training data at scale
Compensation
- Labelbox strives to ensure pay parity across the organization and discuss compensation transparently.
Benefits
- Continuous Growth: Every role requires continuous learning and evolution.
- Direct experience building or contributing to RL environments (Gymnasium/Gym, PettingZoo, or custom environment implementations)
Company info
- Shape the Future of AI
- At Labelbox, we're building the critical infrastructure that powers breakthrough AI models at leading research labs and enterprises.
- Since 2018, we've been pioneering data-centric approaches that are fundamental to AI development, and our work becomes even more essential as AI capabilities expand exponentially.
- About Labelbox
- We're the only company offering three integrated solutions for frontier
- What We're Looking For
Apply directly at Labelbox →Create a free account for alerts like thisView Labelbox immigration profile
This listing is sourced directly from Labelbox's careers page and normalized into a canonical job model.