is home to one of the world's largest collections of human-created content. Scribd’s products leverage one of the world's largest collections of human-created content and intelligent tools to help people move from information access to real understanding and application. This past year, Scribd used Gemini's native PDF understanding and Gemini Enterprise batch prediction to run trust and safety classification across its entire user-generated content corpus of more than 400 million documents, spanning over 12 billion pages, in a matter of months.
Classified 400M+ user-uploaded documents (12B+ pages of text and images) across Scribd and Slideshare Completed the corpus-wide backfill in a matter of months, with Google Cloud scaling batch throughput to meet the timeline Native PDF input meant more than 99% of the corpus was processed as-is, with no OCR, rendering, or screenshotting pipeline to build Gemini Enterprise’s batch prediction at a 50% discount to interactive pricing made LLM classification viable at corpus scale Scribd, Inc. is the parent company to four distinct products: Scribd, Slideshare, Everand, and Fable.
Across Scribd and Slideshare, hundreds of millions of user-uploaded PDFs, presentations, and documents help people find information, build understanding, and finish projects. With that scale comes responsibility. We aim to balance access with protecting our communities. We leverage a mix of human and automated methods to review and best ensure the content on our platforms complies with our community rules. As the corpus continues to grow and technology evolves, this challenge requires even more resources. Understanding a document requires reading its text and its images together, in context.
Classification has to work across all possible use cases, all possible languages, all possible contexts. There is no single solution that can translate cleanly across all of it. And each policy area traditionally demanded its own specialized detection model, which meant either years of in-house engineering effort or specialized vendor solutions that don't fit the economics of a 400-million-document backfill. The team evaluated several off-the-shelf moderation tools and open models, but none delivered the quality they needed at their scale.
