遇见数据集

PluG: A Corpus of 19th- and early 20th-century Ukrainian Texts

收藏
Zenodo2026-06-17 更新2026-06-05 收录
官方服务:

资源简介:

The PluG (Pluperfect GRAC) corpus is a collection of Ukrainian texts from the General Regionally Annotated Corpus of Ukrainian (GRAC: uacorpus.org). It covers texts from 1816 to 1954, including various types such as fiction, news articles, and other writings. The corpus focuses on works from before the mid-20th century and contains texts by 7,590 unique authors and 44 unique translators. The corpus features 42,000 files with 58,676,313 tokens (109M Gemma). It consists of copyright-free classic literature and other old texts suitable for LLM training, computational linguistics studies and education. The texts of the corpus were extracted from printed sources using OCR and corrected manually. It includes some texts written in old orthographical systems (Kulishivka, Zhelekhivka, Skrypnykivka). The texts come from various regions of Ukraine, with many from cities like Kyiv, Lviv, and Kharkiv. PluG includes both original Ukrainian works and translations from other languages. PluG2 is an expanded version of the PluG corpus that contains a larger collection of Western Ukrainian texts from the 1880s to the 1920s written using the orthography system of the time (Zhelekhivka). PluG2 features 73,900,596 tokens. The added texts represent not only a distinctive orthographic system, but also a separate historical variant of literary Ukrainian, which has numerous peculiar grammatical and lexical features and can cause complications when training models oriented to the modern standard. The corpus is available under CC-BY license. It is designed as a dataset for applied linguistic studies, providing a valuable resource for research on Ukrainian literature, language development, and cultural history of the 19th and early 20th centuries. The corpus provides a wide range of metadata for each text, including information about authors, translators, years of publication, genres, styles, and locations. Full tagset used in the meta-annotation are available on the GRAC website: https://uacorpus.org/rozmitka-tekstiv/stili-tematika-i-zhanri It's planned to be updated yearly to keep the resource up-to-date and valuable for researchers. Acknowledgements: A large part of the collection was sourced from open digital libraries, most notably the collection of Western Ukrainian newspapers assembled by Orest Drul (https://zbruc.eu/). We are grateful to Orest Drul, Maksym Bystrytskyi, Mykola Zharkykh, Mykhailo Nazarenko, Nataliia Mykhailivska, and all those who create and maintain open digital libraries.

PluG(超完美GRAC)语料库源自乌克兰通用区域标注语料库(General Regionally Annotated Corpus of Ukrainian,简称GRAC,官网:uacorpus.org),是乌克兰语文本的集合。该语料库涵盖1816年至1954年的文本,类型包括小说、新闻报道及其他各类文稿,核心聚焦于20世纪中期之前的作品,收录了7590位独立作者与44位独立译者的文本。 该语料库包含42000个文件,共计58,676,313个词元(Token),对应109M Gemma模型的训练数据量级。其内容涵盖无版权经典文学作品与其他老旧文本,适用于大语言模型(Large Language Model,简称LLM)训练、计算语言学研究与教学场景。语料库内的文本均通过光学字符识别(Optical Character Recognition,简称OCR)技术从印刷源文件中提取,并经人工校对修正,其中包含部分使用旧正字法体系(库利什夫卡正字法、热列赫夫卡正字法、斯克里普尼克夫卡正字法)撰写的文本。文本来源覆盖乌克兰多个地区,大量文本源自基辅、利沃夫、哈尔科夫等城市。PluG同时收录原创乌克兰语作品与外译乌克兰语译作。 PluG2是PluG语料库的扩展版本,收录了更多1880年代至1920年代的乌克兰西部文本,这些文本均采用当时的热列赫夫卡正字法体系撰写,共计73,900,596个词元。新增文本不仅对应独特的正字法体系,同时代表了文学乌克兰语的独立历史变体——该变体拥有大量独特的语法与词汇特征,在面向现代标准乌克兰语的模型训练过程中可能带来额外复杂度。 该语料库采用CC-BY许可协议发布,专为应用语言学研究打造,为19世纪至20世纪初的乌克兰文学、语言发展及文化历史研究提供了宝贵资源。语料库为每篇文本提供了丰富的元数据,涵盖作者、译者、出版年份、体裁、风格与来源地等信息。元标注所使用的完整标记集可在GRAC官网查询:https://uacorpus.org/rozmitka-tekstiv/stili-tematika-i-zhanri。 该语料库计划每年更新一次,以维持其时效性与对研究者的学术价值。 致谢:本语料库的大量内容源自开放数字图书馆,其中最核心的来源是奥列斯特·德鲁尔(Orest Drul)整理的乌克兰西部报纸馆藏(https://zbruc.eu/)。我们谨向奥列斯特·德鲁尔、马克西姆·比斯特里茨基、米科拉·扎尔基赫、米哈伊洛·纳扎连科、纳塔利娅·米哈伊利夫斯卡以及所有参与创建和维护开放数字图书馆的人士致以诚挚谢意。

提供机构:
Zenodo
创建时间:
2026-04-09
二维码
社区交流群
二维码
科研交流群
商业服务