When it comes to Chinese AI labs using distillation techniques to extract knowledge from frontier model makers, Y Combinator CEO Garry Tan is hoping regulators stay out of it. AI labs should perhaps play the same game. “I would do nothing,” he told CNBC in an interview earlier this week . “We could argue that there should be an American distillation regime. ” He elaborated to TechCrunch that this means he wants smaller, American open-weight AI labs to use the same kind of training techniques on American frontier AI labs, giving the U. a more robust set of open-weight options that aren’t Chinese.

Distillation is when a model maker extensively prompts another model in order to learn how it works and reasons. It is commonly, and legitimately, used by AI labs to help train new models. Anthropic this week released its second report alleging that Chinese labs are engaged in "illicit distillation attacks," hiding their identities to distill without permission and relying on fraud and stolen credentials to do so. Anthropic CEO Dario Amodei had previously publicly called on U. regulators to crack down on distillation.

It's notable that the commander of Silicon Valley's prestigious and prolific startup accelerator doesn't agree. To be clear, Tan isn't advocating for American AI labs to use stolen credentials to distill. He wants them to be free to come in the front door. In fact, his argument is twofold. He feels it's an overreach for AI labs to dictate what their customers can do with the information their models share with them. He also notes that the proprietary AI labs didn't ask permission when they vacuumed up as much human knowledge as they could to train their models.

They famously ingested plenty of copyrighted material without the permission of those intellectual property holders . "Controlling what users and customers do with API calls to closed weight models feels constraining, and there's a role government can play here to normalize the fact that access to intelligence that was trained on broad public access data should itself also be more a form of a public good than something locked away behind restrictive terms of service," he told TechCrunch when asked why American labs should be free to distill, too.