Viral Culture in Early Nineteenth-Century Europe newspaper dataset
收藏资源简介:
Dataset produced during the project Viral Culture in Early Nineteenth-Century Europe. The project traced text reuse by analysing large OCR'd newspaper collections using a BLAST based algorithm. This algorithm produces text clusters. This dataset contains two produced cluster datasets based on two different data collections. For the first dataset, the Austrian ANNO newspaper collection, this dataset contains metadata describing the used newspapers. For the second dataset, German-language newspapers in the Europeana collection, this dataset contains project produced metadata describing the newspapers used by the project, as well as the OCR's content for these newspaper issues. The OCR is produced with Tesseract OCR from digital page images downloaded from the Europeana services.
本数据集源自《十九世纪早期欧洲的病毒式文化传播》(Viral Culture in Early Nineteenth-Century Europe)研究项目。该项目借助基于BLAST的算法,对经光学字符识别(Optical Character Recognition, OCR)处理的大规模报纸馆藏开展文本复用分析,以此追溯文本复用行为。该算法可生成文本聚类簇。本数据集包含基于两套不同馆藏生成的两类聚类数据集:第一类数据集对应奥地利ANNO报纸馆藏,其中附带了本次研究使用的报纸的元数据;第二类数据集对应Europeana馆藏中的德语报纸,其中不仅包含项目生成的、用于描述研究所用报纸的元数据,还涵盖了对应报纸各期的OCR文本内容。上述OCR文本通过Tesseract OCR工具,对从Europeana平台下载的数字化页面图像进行识别后生成。



