MU-NLPC/czech_korpus_kala
收藏资源简介:
该数据集是在专注于捷克文本自动作者识别的硕士论文框架内创建的。它包含从公开可用的在线资源获取的捷克新闻文本,并为以下任务的实验而准备:作者归属(authorship attribution)、作者验证(authorship verification)和按作者聚类(authorship clustering)。数据来源于五个捷克新闻门户网站:Echo24、Stisk online、Deník Referendum、Deník.cz 和 Aktuálně.cz。数据集规模包括5个来源门户、524位作者和17,213份文档,每位作者最少15份、最多50份文档。数据经过预处理,过滤了短于1800字符或300单词的文本,最终保留至少1500字符或250单词的文本,并进行了清理和标准化。数据集主要用于风格测量、作者识别、文本分类和捷克语自然语言处理研究。限制包括每位作者仅来自一个来源门户,且仅包含新闻文本。数据集基于公开文本创建,使用时需尊重原始来源条款和版权。
This dataset was created as part of a masters thesis focused on automatic authorship recognition of Czech texts. It contains Czech journalistic texts obtained from publicly available online sources and prepared for experiments in tasks: authorship attribution, authorship verification, and authorship clustering. Data sources include five Czech news portals: Echo24, Stisk online, Deník Referendum, Deník.cz, and Aktuálně.cz. The dataset scale comprises 5 source portals, 524 authors, and 17,213 documents, with a minimum of 15 and maximum of 50 documents per author. Preprocessing involved filtering texts shorter than 1800 characters or 300 words, with final retention of texts meeting at least 1500 characters or 250 words, followed by cleaning and normalization. The dataset is intended for research in stylometry, authorship recognition, text classification, and Czech natural language processing. Limitations include each author being from only one source portal and the dataset containing only journalistic texts. It is based on publicly available texts, and use must respect original source terms and copyrights.




