Bespokelabs

Bespokelabs

Backend Engineer

Mountain View

Sponsorship not specifiedDetected 56 days ago
Node.jsAWSGCPCloud PlatformsFirewallResearchCollaboration

About the role

  • This is a hard systems problem disguised as an AI job.
  • As the tasks agents can complete keep lengthening, the environments that train them have to stay coherent across far longer horizons than anything that exists today.
  • That means sandboxing and isolation you can trust, execution that's fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot.

Responsibilities

  • Design and own the sandboxing and execution layer that environments run inside.
  • Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.
  • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known-good state, patch, and continue.
  • Own the performance characteristics of the platform: throughput, latency, and cost-per-rollout at scale.
  • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.
  • Build the observability that lets us understand what's happening inside thousands of concurrent, long-running rollouts.
  • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.
  • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.

Requirements

  • Ability to use modern tools such as Claude Code effectively.
  • Ability to translate between research needs and infrastructure requirements.

Nice to have

  • Experience with RL training or evaluation infrastructure, or the execution layer for agent rollouts.
  • Experience with checkpoint/snapshot-restore systems, CRIU, or distributed state management.
  • Background in high-throughput, low-latency execution systems.
  • Contributions to widely-used infrastructure, datasets, benchmarks, or open-source systems.
  • Previous experience in a research engineering or infrastructure role at an AI or systems-heavy company.
  • Experience making systems fast and cheap - profiling, scheduling, resource utilization, and cost optimization at scale.
  • Proficiency with cloud platforms (GCP, AWS) and distributed computing.
  • Strong engineering fundamentals and a systematic approach to testing, validation, and reliability.

Skills

  • Extend execution to long-horizon and multi-node environments, where an agent operates across many tools and services over hours or days.
  • Performance & Scale
  • Profile and remove bottlenecks across the stack, from container startup to environment teardown.
  • Environment Platform
  • Collaboration & Production Excellence
  • Scale prototypes into production systems with reproducible workflows and high engineering standards.

Compensation

  • Competitive salary and equity

Benefits

  • Health coverage, and the opportunity to work directly with the world's leading AI research labs

Company info

  • You'll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

This listing is sourced directly from Bespokelabs's careers page and normalized into a canonical job model.