遇见数据集

LibMovIt: text corpus of travel literature

收藏
Zenodo2025-10-15 更新2026-05-26 收录
官方服务:

资源简介:

LibMovIt: text corpus of travel literature is a textual resource created within the LibMovIt project (Progetto finanziato dall’Unione Europea – NextGenerationEU a valere sul Piano Nazionale di Ripresa e Resilienza (PNRR) – Missione 4 Istruzione e ricerca – Componente 2 Dalla ricerca all’impresa – Investimento 1.1, Avviso Prin 2022 indetto con DD N. 104 del 2/2/2022, Progetto dal titolo LIBMOVIT – Libraries on the move: scholars, books, ideas trave-ling in Italy in the 18th century, codice proposta 2022CP88KY). Version 1.0 of the corpus contains 52 works for 7,9 milion words: 27 in English (3,590,000 words), 12 in French (2,065,000), 8 in German (1,550,000), 4 in Italian (450,000) and 1 in Spanish (255,000). A detailed description of the corpus and the status of each text (tags: "Revision completed" or "Revision to be completed") is available in a Zotero library at the following link: https://www.zotero.org/groups/5540957/libmovit/library Texts are published in .txt format. They are the result of both automatic text recognition and acquisition from other projects (indicated as tags in the corpus description). All the newly recognised texts have been reviewed with scripts to clean up the most common errors and delete paratextual elements (page numbers, catchwords, signature marks etc.). The editors of the corpus made manual corrections in all the texts, however, due to their lenghth, some of them still need further revision. For this reason, minor updates of the corpus (1.1, 1.2, 1.3 etc.) will be released regularly to improve the texts with further corrections, text mark-up and conversion to other formats; a major update of the corpus will be released once a year and will also include new texts. Additional information about the corpus development are described in the papers listed in the references section.

LibMovIt:旅行文学文本语料库是LibMovIt项目所创建的文本资源。该项目由欧盟资助,依托下一代欧盟(NextGenerationEU)框架,基于《国家复苏与韧性计划》(Piano Nazionale di Ripresa e Resilienza, PNRR)——任务4:教育与研究——组件2:从研究到企业——投资1.1,对应2022年PRIN科研公告,该公告通过2022年2月2日第104号行政令发布,项目全称为《LIBMOVIT——移动中的图书馆:18世纪意大利的学者、书籍与思想传播》,提案编号为2022CP88KY。 该语料库1.0版本共收录52部作品,总计790万字:其中英语作品27部(359万字)、法语作品12部(206.5万字)、德语作品8部(155万字)、意大利语作品4部(45万字),西班牙语作品1部(25.5万字)。语料库的详细说明及每篇文本的状态(标签为"已完成修订"或"待完成修订")可通过以下链接访问的Zotero文库获取:https://www.zotero.org/groups/5540957/libmovit/library 所有文本均以纯文本(.txt)格式发布,其来源涵盖自动文本识别成果及其他项目的授权获取内容(语料库说明中以标签标注来源)。所有新识别的文本均通过脚本清理常见错误,并删除副文本元素,包括页码、页脚衔接词、签名标记等。此外,语料库编辑人员已对全部文本开展手动校正,但由于部分文本篇幅较长,仍需开展进一步修订。 为此,本语料库将定期发布小版本更新(如1.1、1.2、1.3等),通过进一步校正、文本标注及格式转换优化文本内容;每年将发布一次大版本更新,同时新增收录文本。 有关语料库开发的更多详细信息,请参阅参考文献部分列出的相关学术论文。

提供机构:
CNR-ILIESI
创建时间:
2025-10-15
二维码
社区交流群
二维码
科研交流群
商业服务