burmese-synthetic-speech-corpus
收藏资源简介:
该数据集是一个高质量、人工标注的缅甸语文本集合,专门为二元文本分类任务设计。它包含1000个缅甸语文本条目,在垃圾邮件和非垃圾邮件两个类别上完全平衡,各含500个样本。数据集旨在促进缅甸语垃圾邮件过滤模型的开发与评估,数据来源于多种数字平台,包括互联网评论、社交媒体互动以及专业和个人电子邮件通信。所有文本都经过清理和标准化处理,使用有效的缅甸Unicode编码,确保与现代自然语言处理工具包的兼容性。标注过程采用严格的同行评审机制:两位创建者独立标注整个语料库,随后通过共识建立阶段逐一验证每个标签的准确性和一致性。数据集以CSV格式提供,包含text和label两列。需要注意的是,标注结果受到标注者主观视角的影响,反映了创建者对缅甸数字环境中垃圾邮件与非垃圾邮件的解读,且数据集捕捉的是创建时的语言趋势和垃圾邮件模式快照。
This dataset is a high-quality, manually annotated Burmese text collection specifically designed for binary text classification tasks. It comprises 1000 Burmese text entries, with a perfectly balanced distribution across the two categories of spam and non-spam, containing 500 samples for each category. This dataset is intended to support the development and evaluation of Burmese spam filtering models, with data sourced from multiple digital platforms including internet comments, social media interactions, as well as professional and personal email communications. All texts have been cleaned and standardized using valid Burmese Unicode encoding to ensure compatibility with modern natural language processing toolkits. The annotation process follows a strict peer review workflow: two creators independently annotated the entire corpus, followed by a consensus-building phase to verify the accuracy and consistency of every individual label. The dataset is distributed in CSV format, featuring two columns: text and label. It is important to note that the annotation results are subject to the subjective perspectives of the annotators, reflecting the creators' understanding of spam and non-spam within the Burmese digital ecosystem, and the dataset captures a snapshot of linguistic trends and spam patterns at the time of its construction.




