Institutional Books — Visual Elements
收藏资源简介:
Institutional Books — Visual Elements 数据集由哈佛法学院图书馆等机构创建,收录了从哈佛图书馆Google Books数字化藏书中提取的2260万视觉元素,涵盖插图、照片、版画等历史文献内容,并附带766,992,447个o200k_base文本标记的实验性标题。该数据集通过开源pipeline对近百万卷图书的页面扫描进行检测、分类、去重和标题生成,旨在弥补视觉语言模型对历史视觉内容理解不足的缺陷,为人工智能模型训练与数字人文研究提供高质量、多样化的历史视觉语料。
The Institutional Books — Visual Elements dataset was developed by institutions including Harvard Law School Library. It houses 22.6 million visual elements extracted from the digitized Google Books holdings of Harvard Library, encompassing historical documentary content such as illustrations, photographs, and prints, and includes experimental titles associated with 766,992,447 o200k_base text tokens. This dataset utilizes an open-source pipeline to detect, classify, deduplicate, and generate titles for page scans of nearly one million book volumes, aiming to address the gap in visual language models' comprehension of historical visual content, and provide high-quality, diverse historical visual corpora for artificial intelligence model training and digital humanities research.




