遇见数据集

European Literary Text Collection (ELTeC)

收藏
SSH Open MarketPlace2024-09-30 更新2024-10-05 收录
官方服务:

资源简介:

The European Literary Text Collection (ELTeC) is a multilingual collection of corpora of literary texts that are comparable in nature, scope and quality across several European languages. Its availability is an essential condition for the creation, evaluation and use of multilingual tools and methods of analysis for literary texts. Novels have been chosen among major literary genres for availability and size. Chronological limits are due to constraints related to copyright and availability of quality full texts. – This release contains 14 collections of at least 50 novels, 8 of which have reached 100 novels, and a total of over 1200 novels. Collections in this release include: Czech, German, English, French, Hungarian, Norwegian, Polish, Portuguese, Romanian, Slovenian, Spanish, Serbian, Swedish and Ukrainian. Note that this is an umbrella release, referencing the collections included in the release, but not containing the actual files. – ELTeC is one of the key outcomes of the COST Action ‘Distant Reading for European Literary History’ (CA16204) that ran from 2017 to 2022.

欧洲文学文本数据集(European Literary Text Collection,ELTeC)是一个多语言文学文本语料库集合,在多种欧洲语言间具备性质、规模与质量上的可比性。该数据集的可获取性,是开发、评估与应用面向文学文本的多语言分析工具与分析方法的必要前提。本次数据集选取了主流文学体裁中的小说类别,以保障语料的可获取性与规模体量。数据集的时间范围限制,源于版权约束与高质量完整文本的可获取性限制。—— 本次发布版本包含14个语料库集合,每个集合至少包含50部小说,其中8个集合的小说数量已达100部,整体总规模超过1200部小说。本次发布的语料库集合涵盖:捷克语、德语、英语、法语、匈牙利语、挪威语、波兰语、葡萄牙语、罗马尼亚语、斯洛文尼亚语、西班牙语、塞尔维亚语、瑞典语以及乌克兰语。请注意,本次发布为总览式发布,仅列明本次发布所包含的语料库集合,并未附带实际的文本文件。—— ELTeC是2017至2022年实施的COST行动‘面向欧洲文学史的远读’(Distant Reading for European Literary History,CA16204)的核心成果之一。

创建时间:
2024-09-30
二维码
社区交流群
二维码
科研交流群
商业服务