Scribd, Inc. classifies millions of documents on Gemini Enterprise
Scribd, Inc. is home to one of the world's largest collections of human-created content.
Scribd’s products leverage one of the world's largest collections of human-created content and intelligent tools to help people move from information access to real understanding and application.
This past year, Scribd used Gemini's native PDF understanding and Gemini Enterprise batch prediction to run trust and safety classification across its entire user-generated content corpus of more than 400 million documents, spanning over 12 billion pages, in a matter of months.
Here were the results:
- Classified 400M+ user-uploaded documents (12B+ pages of text and images) across Scribd and Slideshare
- Completed the corpus-wide backfill in a matter of months, with Google Cloud scaling batch throughput to meet the timeline
- Native PDF input meant more than 99% of the corpus was processed as-is, with no OCR, rendering, or screenshotting pipeline to build
- Gemini Enterprise’s batch prediction at a 50%...
Copyright of this story solely belongs to cloud.google.com. To see the full text click HERE