We hereby declare September to be scalability month! As the world prepares for a surge of agentic fleets, we are shoring up our AI infrastructure and orchestration offerings to gracefully — and quickly — respond to that demand, all while maintaining workload isolation and security, and keeping costs in check. Read on to learn how these enhancements manifest across Google Cloud’s compute, network, storage, and orchestration offerings, plus new ways customers are using Google Cloud AI infrastructure, and third-party industry validation of our strategy.
Google Kubernetes Engine updates: The GKE team is all about improving the scalability of the platform, and in September, those improvements came in many shapes and sizes: New feature: Need an execution runtime with higher density for your agentic workloads? We engineered the new open-source GKE Agent Substrate to run millions of sandboxes with 10x higher density than standard container runtimes. Agent Substrate also delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation.
Product update: GKE now has scale-to-zero capabilities built-in. No need to configure complex components to scale your workloads down, thanks to the HPA with the Autoscaling Metric and support for KEP-2021, which do the job for you, out of the box. Read the blog to learn more. Product update: Further, the GKE HPA (with the above-mentioned Autoscaling Metric) now lets you scale up and down based on custom PromQL metrics, in addition to standard metrics, allowing you to trigger workloads according to conditions that are meaningful and unique to your business.
