遇见数据集

OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)

收藏
Zenodo2020-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945. At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012). corpus-entropy.pkl entropy rate per document page corpus-language.pkl language per document page corpus.zip fulltext corpus (extracts to .txt format) de_corpus.zip German sub-corpus (extracts to .txt format) selection_de.pkl Selection list of German documents xml2csv_alto.csv fulltext corpus per document page (incl.OCR word confidences) <em>Sources</em> Marco Lui and Timothy Baldwin. 2012. Langid.py: An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations, ACL ’12, pages 25–30, Stroudsburg, PA, USA. Association for Computational Linguistics

SBB的数字馆藏涵盖1470年至1945年间的153,942件数字化作品。 发布时,已有28,909件作品完成了光学字符识别(Optical Character Recognition,OCR)处理,共生成4,988,099页全文。对于每一段包含OCR文本的页面,其语言均通过langid工具(Lui与Baldwin,2012)完成判定。 corpus-entropy.pkl:各文档页面的熵率 corpus-language.pkl:各文档页面的语言标注信息 corpus.zip:全文语料库(解压后为.txt格式) de_corpus.zip:德语子语料库(解压后为.txt格式) selection_de.pkl:德语文档精选列表 xml2csv_alto.csv:各文档页面的全文语料(包含OCR词置信度) 参考文献 Marco Lui与Timothy Baldwin,2012年。Langid.py:一款现成的通用语言识别工具。收录于《ACL 2012系统演示会议论文集》(ACL ’12),美国宾夕法尼亚州斯特劳兹堡,计算语言学协会,第25–30页。

提供机构:
Zenodo
创建时间:
2019-06-26
二维码
社区交流群
二维码
科研交流群
商业服务