Bosnian Corpus (v1.0): Cleaned Web and SMS Text for Entropy and NLP Research
收藏资源简介:
This record provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative analysis of language entropy, "language energy" and modern NLP tasks. The corpus is built from three publicly available resources released in the CLARIN.SI repository:(1) The Sarajevo Corpus of SMS Messages in Bosnian 1.1,(2) Bosnian web corpus bsWaC 1.1, and(3) Bosnian web corpus CLASSLA-web.bs 1.0.All sources were converted to plain text, cleaned, normalised, partially deduplicated, and merged into a single consistent dataset. The final corpus contains approximately 6.18 GB of text (≈ 6,182,905,888 bytes), 46,258,935 lines and 942,515,845 tokens.The web portion is organised into several “super-genres” (News, Opinion, Forum/Chat, Info/HowTo, Legal/Admin, Literature, Ads/Promo, Mix/Other).For each super-genre a separate text file is provided, together with one global file that concatenates all genres for entropy estimation and language-model training. Cleaning focuses on removing technical noise that would bias frequency distributions and entropy estimates, while preserving the linguistic signal:– Unicode normalisation (UTF-8, NFC),– correction of common mojibake artefacts,– removal of URLs, e-mail addresses, file names, boilerplate and CMS/navigation lines,– filtering of lines with a high proportion of non-letter characters,– optional digit normalisation and lowercasing,– language filtering to keep primarily Bosnian text. Files in this record:– bosnian_corpus_all.txt (full corpus, all genres),– per-genre text files (news, forum, opinion, info/howto, legal/admin, literature, ads/promo, mix/other),– README.txt with dataset description,– two accompanying research papers (Bosnian and English), uploaded separately as PDF files. Code availability:Preprocessing, cleaning and entropy-calculation scripts are publicly available on GitHub:https://github.com/H4sK0/bosnian-corpus-pipeline Licence:This corpus is released under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) licence.Users must credit this Zenodo record and the original source corpora (Sarajevo SMS 1.1, bsWaC 1.1, CLASSLA-web.bs 1.0), and must distribute derivative corpora under the same or a compatible licence. Suggested citation:Hasan Kahrimanović (2025). Bosnian Corpus (v1.0): Cleaned Web and SMS Text for Entropy and NLP Research. Zenodo. DOI: [assigned by Zenodo].
本数据集提供了经过清洗且标注了体裁的当代波斯尼亚语语料库(corpus),专为语言熵、“语言能量”的量化分析以及现代自然语言处理(Natural Language Processing)任务设计。 该语料库源自CLARIN.SI存储库中发布的三份公开可用资源:(1) 波斯尼亚语短信萨拉热窝语料库(Sarajevo Corpus of SMS Messages in Bosnian)1.1版,(2) 波斯尼亚语网页语料库bsWaC 1.1版,以及(3) 波斯尼亚语网页语料库CLASSLA-web.bs 1.0版。所有源数据均被转换为纯文本格式,经过清洗、归一化、部分去重处理后,合并为一个统一的数据集。 最终语料库包含约6.18 GB的文本数据(约6,182,905,888字节),共计46,258,935行文本与942,515,845个词元(Token)。该语料库的网页部分被划分为多个“超体裁(super-genres)”,涵盖News、Opinion、Forum/Chat、Info/HowTo、Legal/Admin、Literature、Ads/Promo、Mix/Other。针对每个超体裁均提供了独立的文本文件,同时附带一个整合了所有体裁的全局文件,用于语言熵估算与语言模型训练。 清洗流程旨在移除会干扰频率分布与语言熵估算的技术噪声,同时保留语言信号: – Unicode归一化(UTF-8,NFC), – 修复常见的乱码痕迹(mojibake artefacts), – 移除URL、电子邮箱地址、文件名、样板文本(boilerplate)以及内容管理系统(CMS)导航栏文本, – 过滤非字母字符占比过高的行, – 可选的数字归一化与小写转换, – 语言过滤,仅保留以波斯尼亚语为主的文本。 本数据集中包含以下文件: – bosnian_corpus_all.txt(完整语料库,包含所有体裁), – 各体裁独立文本文件(对应新闻、论坛、评论、信息/操作指南、法律/行政、文学、广告/推广、混合/其他), – README.txt(数据集说明文档), – 两篇配套研究论文(分别为波斯尼亚语与英语版本),以PDF文件形式单独上传。 代码获取:预处理、清洗与语言熵计算脚本已公开托管于GitHub平台:https://github.com/H4sK0/bosnian-corpus-pipeline 许可协议:本语料库采用知识共享署名-相同方式共享4.0国际许可协议(Creative Commons Attribution-ShareAlike 4.0 International, CC BY-SA 4.0)进行发布。使用者需注明本Zenodo数据集以及原始源语料库(萨拉热窝短信语料库1.1版、bsWaC 1.1版、CLASSLA-web.bs 1.0版),且衍生语料库需采用相同或兼容的许可协议进行分发。 推荐引用格式:Hasan Kahrimanović(2025)。波斯尼亚语语料库(v1.0):用于语言熵与自然语言处理研究的清洗版网页及短信文本。Zenodo。DOI:[由Zenodo分配]。



