Modern reinforcement learning (RL) systems rely on thousands of rollout environments running in parallel. For large language model (LLM) training using Verifiable Rewards (RLVR) each rollout environment runs in a tightly isolated sandbox so state never leaks between environments. The faster you can recycle sandboxes, the more training steps you can run per hour. That means higher rollout density per host which directly lowers your cost per step. But density is a tuning knob with two failure modes. Push it too high and you hit a tail latency cliff.

Play it too safe and you overprovision, leaving money on the table. The sweet spot depends on your instance type, workload profile, and how your training algorithm amplifies stragglers. This post shows how to find that setting. We walk through the tunables that control density on Amazon EC2 metal. We then introduce labsweep , a measurement harness for mapping your latency and density curve. Throughout the post, we draw on lessons learned from production deployments with customers. The post is written for readers familiar with Kernel-based Virtual Machine (KVM), Firecracker, and container orchestration.

The AWS Nitro System offloads networking, storage, and security to dedicated hardware. Amazon EC2 instances expose a stable CPU topology you can inspect and pin workloads against. On Amazon EC2 metal, applications run directly on the underlying hardware. Firecracker is a lightweight virtual machine monitor (VMM) designed for high density workloads. Each microVM boots quickly, runs its own Linux kernel, and has a small memory footprint. Running Firecracker on Amazon EC2 metal combines direct access to underlying hardware with lightweight virtualization to support high density sandboxing.

The key question is how far you can increase density before tail latency hurts the RL training loop. The following diagram shows the reference architecture of a production RL sandbox node. This architecture serves as the foundation for the rest of the post. An Amazon EC2 metal instance runs a host agent that manages Firecracker microVMs, CPU pinning, non-uniform memory access (NUMA)-aware memory placement, and the warm pool used for dispatch. The host agent runs as a DaemonSet on each Amazon EC2 metal node in the Kubernetes cluster.