Everyone's waiting for Nvidia to confirm this week's most interesting tech deal: A reported $13 billion acquisition of Hugging Face, a platform for sharing open-weight AI models and benchmarks. Now best known as the target for a team of reward-hacking OpenAI agents, Hugging Face is at the center of the ecosystem of developers building and deploying LLMs that aren't owned by frontier labs. Think of it as a kind of GitHub for the AI era. Rumors of that deal come after Nvidia struck a $6 billion agreement with Poolside, an open-weight model builder, that will see most of its employees move to the chip-making giant.
And two weeks ago, Stripe acquired OpenRouter, the top provider of open-weight models to businesses, for more than $7 billion. That's a lot of capital pouring into a sector based on giving stuff away, and it reflects the latest trends in the AI sector. For Nvidia, there's a need to avoid further dependence on its deals with the major hyperscalers and frontier labs. That's particularly the case when major AI model builders like OpenAI and Google are also building their own inference chips, like OpenAI's Jalapeño , whose capabilities were announced this week.
If model builders are making chips, Nvidia wants a chunk of the model-making business. Nvidia already builds its own Nemotron family of open-weight models, but their uptake hasn't been huge. By taking control of the largest U. developer space for open models, the company will have access to a mass of users it can drive to its chips and standards. There are also growing questions about the cost of AI inference, which has companies exploring cheaper models built by Chinese companies like Moonshot, DeepSeek, and Alibaba.
Right now, adoption is relatively small but growing — just 6% of companies use open-weight models, according to a survey of spending data by Ramp, or just 2% of software engineers surveyed by Jellyfish , which makes tools for developers. Nik Albarran, the AI product lead at Jellyfish, told TechCrunch that open-weight models are primarily used by companies whose products rely on repeated inference workloads, like those providing customer service chats. Because these are high-volume tasks with a lot of repetition, an open-weight model can be tuned to answer the questions cheaply.
