As customers modernize to lakehouse architectures, they are standardizing on open formats such as Apache Iceberg to create a shared data estate across compatible engines. This enables you to build AI-native lakehouses that turn data to semantic knowledge, enable proactive action, and operate at an agentic scale. One of the key components of a lakehouse is the catalog, and in the Apache Iceberg environment, that usually means the Iceberg REST Catalog. An Apache Iceberg catalog is responsible for maintaining table pointers, handling atomic commits, and serving as the single source of truth for table locations.
But as large enterprise organizations modernize to lakehouses, they have begun to realize that they need a highly scalable and available managed catalog as a part of their lakehouse architecture. This becomes even more important as querying scales with agents. To support a large amount of repeated small queries from agents, you will need to build on top of a managed catalog that provides atomicity, consistency, availability and concurrency at massive scale.
In this blog, we explore the challenges a managed catalog faces in modern cloud environments at agent scale, and show you how Google Cloud’s serverless Lakehouse runtime catalog can help address them. Powered by Spanner , Google Cloud’s always-on database with virtually unlimited scale, and built to meet the open Apache Iceberg REST catalog specification, the Lakehouse runtime catalog is the highly scalable and available foundation you need for the agentic era.
When speaking with data engineers and infrastructure leads running production analytics at scale in a Lakehouse, the following core pain points consistently emerge and need to be solved by a managed catalog: Atomic commits and concurrency control: Iceberg guarantees ACID transactions via optimistic concurrency control (OCC). A managed catalog must implement a bulletproof atomic compare-and-swap (CAS) operation to swap the current metadata pointer. High availability and operational maintenance: Because queries fail immediately if the catalog is down, a managed catalog becomes a critical Tier-1 service.
