Public Domain 12M
收藏资源简介:
Public Domain 12M(PD12M)是由Spawning创建的大规模图像-文本数据集,包含1240万张高质量的公共领域及CC0许可图片,搭配合成字幕,旨在训练文本到图像的模型。该数据集是目前最大的公共领域图像-文本数据集,以其庞大的规模和明确的版权声明,为AI模型的训练提供了坚实的基础,同时最小化了版权担忧。PD12M的数据来源包括画廊、图书馆、档案馆、博物馆(GLAM)以及Wikimedia Commons等,通过精心筛选和治理,确保了数据的质量和安全性。数据集的构建过程涵盖了从图像收集、版权验证、图像下载、内容过滤到字幕生成等多个步骤。特别地,PD12M通过Source.Plus平台引入了社区驱动的数据治理机制,以支持数据集的持续改进和维护。该数据集不仅为AI领域提供了丰富的训练资源,也为负责任的AI实践提供了范例,促进了公共AI资源的保护和利用。
Public Domain 12M (PD12M) is a large-scale image-text dataset created by Spawning. It contains 12.4 million high-quality public domain and CC0-licensed images paired with synthetic captions, and is developed for training text-to-image models. As the largest public domain image-text dataset to date, PD12M relies on its massive scale and clear copyright statements to provide a solid foundation for AI model training while minimizing copyright-related concerns. The dataset sources include galleries, libraries, archives, museums (GLAM), Wikimedia Commons and other similar platforms. Through rigorous screening and governance, the quality and safety of the dataset are guaranteed. Its construction process covers multiple steps such as image collection, copyright verification, image downloading, content filtering and caption generation. Notably, PD12M introduces a community-driven data governance mechanism via the Source.Plus platform to support the continuous improvement and maintenance of the dataset. This dataset not only provides abundant training resources for the AI field, but also serves as a paradigm for responsible AI practices, promoting the protection and utilization of public AI resources.




