When scaling up agentic reinforcement learning (RL) and evaluation across massive parallel rollouts, frontier AI labs inevitably hit a bottleneck: Expensive GPU clusters sit idle, waiting minutes for CPU sandbox cold-starts, plus thousands of multi-gigabyte SWE-bench -style image pulls and scheduling backlogs. It’s a sandbox infrastructure problem that silently slows down your research and burns your training budget.

To solve this fundamental infrastructure bottleneck, today we are introducing GKE Agent Sandbox optimized for RL along with the Agent Sandbox RL orchestration SDK , plus native integrations for popular RL gyms and harnesses, now generally available. As the operating system for modern AI, Kubernetes has evolved to power massive GPU/TPU training clusters and distributed inference. Now Kubernetes is expanding to drive the next AI compute frontier: agents. But unlike static workloads, agentic workloads evolve rapidly, so infrastructure must evolve just as fast.

Rather than guessing at what RL researchers needed, we placed Kubernetes itself on an auto-research and verification loop driven by performance benchmarks and evaluations . We used heavy agentic benchmarks like SWE-bench to intentionally stress-test and break our own clusters. Every bottleneck that surfaced — from etcd timeouts to GPU idle spikes — was fed back into our development cycle to refine GKE’s core primitives.

This resulted in a purpose-built sandbox layer for agentic RL and eval workloads that features: With this new primitive, AI labs and agent-native startups can now reliably run large scale agentic RL trajectories and evals simultaneously, minimizing accelerator idle time and drastically accelerating their research velocity. Before we talk about the solution, let’s be precise about what makes agentic RL so demanding for infrastructure in the first place.