遇见数据集

19th Century United States Newspaper Advert images with 'illustrated' or 'non illustrated' labels

收藏
Zenodo2021-11-08 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (chroniclingamerica.loc.gov/). [The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in <em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the Beyond Words crowdsourcing project. source: https://news-navigator.labs.loc.gov/ One of these categories is 'advertisements. This dataset contains a sample of these images with additional labels indicating if the advert is 'illustrated' or 'not illustrated'. The data is organised as follows: The images themselves can be found in `images.zip` `newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset. `ads.csv` contains the labels for the images as a CSV file `sample.csv` contains additional metadata about the images (based on the newspapers those images came from). This dataset was created for use in an under-review Programming Historian tutorial (http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong> The metadata CSV file contains the following columns: - filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url

本数据集包含源自《报纸导航器(Newspaper Navigator)》(news-navigator.labs.loc.gov/)的图像,该数据集提取自美国国会图书馆《编年史美国(Chronicling America)》馆藏(chroniclingamerica.loc.gov/)。<em>《报纸导航器数据集(Newspaper Navigator Dataset)》</em>包含从16,358,041页<em>Chronicling America</em>历史报纸页面中提取的视觉内容。上述视觉内容通过在第一次世界大战时期<em>Chronicling America</em>页面的标注集上训练的目标检测模型(object detection model)识别得到,其中标注工作由志愿者作为"超越文字(Beyond Words)"众包项目的一部分完成。来源:https://news-navigator.labs.loc.gov/。其中一个类别为"广告(advertisements)"。本数据集为该数据集的样本子集,并额外添加了标注,用于指示广告属于"带插图(illustrated)"还是"无插图(not illustrated)"。数据组织形式如下:图像文件均存放于`images.zip`压缩包中;`newspaper-navigator-sample-metadata.csv`包含源自《报纸导航器数据集》的每张图像的元数据;`ads.csv`以CSV格式存储了图像的标注标签;`sample.csv`包含关于图像的额外元数据(基于图像来源报纸的相关信息)。本数据集为一篇正在审稿中的《编程史学家(Programming Historian)》教程(http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1)所创建,其核心目标是提供一个贴合实际的示例数据集,用于教授如何针对数字化遗产材料开展计算机视觉相关工作。本数据集在此分享,以期对其他研究者有所助益。<strong>本数据文档仍在完善中,待《编程史学家》教程正式公开上线后将进行更新。</strong> 元数据CSV文件包含以下列:<br>- filepath(文件路径)<br>- pub_date(出版日期)<br>- page_seq_num(页面序号)<br>- edition_seq_num(版次序号)<br>- batch(批次)<br>- lccn(美国国会图书馆控制号,Library of Congress Control Number)<br>- box(盒号)<br>- score(置信度得分)<br>- ocr(光学字符识别结果,Optical Character Recognition)<br>- place_of_publication(出版地点)<br>- geographic_coverage(地理覆盖范围)<br>- name(报纸名称)<br>- publisher(出版方)<br>- url(资源链接)<br>- page_url(页面链接)<br>- month(月份)<br>- year(年份)<br>- iiif_url(国际图像互操作框架链接,International Image Interoperability Framework URL)

提供机构:
Zenodo
创建时间:
2021-11-08
二维码
社区交流群
二维码
科研交流群
商业服务