Every week, I talk with founders who are building at an unbelievable pace. Teams are moving from inception to product-market fit faster than ever, with foundation models wired deeply into their core product workflows. Yet as startup architectures mature, a clear divide has emerged between teams struggling with margins and those scaling sustainably. The most effective engineering teams have abandoned the one-size-fits-all model strategy. In the early days of LLMs the default architecture was simple: send every interaction to the largest model available.

But as applications move into production, serving millions of people and running autonomous multi-agent workflows, relying on a single frontier model starts to strain in three places: Latency penalties: Relying entirely on cloud round trips makes it difficult to deliver the sub-second responsiveness that interactive mobile and desktop apps require. Infrastructure overhead: Self-hosting large open models with more than 70 billion parameters forces early-stage teams to act like infrastructure providers, pulling senior engineers on cluster provisioning and multi-GPU orchestration.

Margin erosion: Sending high-frequency, structured tasks (like intent routing, JSON extraction, or status validation) to general-purpose frontier endpoints spends capital that could be funding product differentiation. Great engineering teams pick the right tool for each job. Most production requests don’t require a frontier generalist, and routing every call to one can actually slow your product down. Instead, the winning pattern is a compound AI stack: pairing frontier models for complex synthesis with compact, open-weight models that you can tune, control, and run anywhere.

It’s for these reasons that an open model like Gemma belongs in your model lineup. With more than one billion downloads across the developer community , Gemma 4 is our most capable open model family to date, using the same foundational research and technology behind the Gemini models. Gemma is built by Google DeepMind using the same foundational research and architecture advances behind the Gemini models. Because they share common DNA and developer tooling, your team can prototype in Google AI Studio and design hybrid architectures where Gemini and Gemma work together.