MU-NLPC/czech_corpus_authorship_recognition
收藏资源简介:
该数据集是捷克作者识别语料库(Kala),由Adam Kala在硕士论文中创建,专注于捷克文本的自动作者识别。它包含从五个公开可用的捷克新闻门户网站(Echo24、Stisk online、Deník Referendum、Deník.cz、Aktuálně.cz)收集的新闻文本,用于作者归属、作者验证和作者聚类等任务。数据集包括524位作者的17,213个文档,每个作者有15到50个文档,并经过预处理,过滤掉短文本(如少于1800字符或300单词),最终保留至少1500字符或250单词的文本,并进行清理和标准化。数据集主要用于风格测量、作者识别、文本分类和捷克自然语言处理研究,但有限制:每个作者仅来自一个来源,且仅包含新闻文本。使用时应尊重原始来源的条款和版权,仅用于研究目的。
The dataset is the Czech Authorship Recognition Corpus (Kala), created by Adam Kala for a masters thesis focused on automatic authorship recognition of Czech texts. It contains Czech journalistic texts collected from five publicly available online sources (Echo24, Stisk online, Deník Referendum, Deník.cz, Aktuálně.cz), prepared for tasks such as authorship attribution, authorship verification, and authorship clustering. The dataset includes 17,213 documents from 524 authors, with each author having between 15 to 50 documents. Preprocessing involved filtering out short texts (e.g., less than 1800 characters or 300 words), and only texts meeting at least one condition (1500 characters or 250 words) were retained in the final cleaned version, with further cleaning and normalization per source. It is intended for research in stylometry, authorship recognition, text classification, and natural language processing for Czech, with limitations: each author comes from only one source portal, and the dataset consists solely of journalistic texts. Legal and ethical notes require respecting the terms of the original sources and copyright, and it is published for research purposes.




