遇见数据集

PolCSBD :Political Counter Speech BD

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset (PolCSBD) was developed to address a critical gap in natural language processing: the detection of political counter-speech in low-resource, code-mixed languages. Our foundational hypothesis was that counter-speech cannot be accurately classified by looking at a single comment in isolation; it fundamentally requires the context of the preceding statement. Additionally, we hypothesized that social media users in Bangladesh heavily use "Banglish" (a phonetic mix of English alphabets and Bengali vocabulary) alongside native Bengali script, which creates a major barrier for standard text classification models. The dataset provides over 10,000 contextual pairs of social media text extracted from political discussions. Each row is structured as a direct conversation, containing a "parent_text" (the initial statement) and a "reply_text" (the direct response). The data demonstrates the complex linguistic reality of the region, featuring native Bengali script, fully Romanized Bengali, and hybrid sentences. It effectively captures how internet users employ historical references, aggressive debate tactics, and sarcasm to challenge political narratives. How to interpret and use the data: This dataset is heavily optimized and provided in a machine-learning-ready format, making it ideal for researchers looking to train, fine-tune, or benchmark Transformer models (such as mBERT, XLM-RoBERTa, or BanglaBERT). It contains exactly three columns: parent_text: The contextual baseline statement, which has been preprocessed to remove noise. reply_text: The responding statement, similarly preprocessed. label: A binary integer classification. A value of '1' indicates Counter-Speech (the reply actively disputes, corrects, or challenges the parent text with a counter-narrative). A value of '0' indicates Non-Counter Speech (the reply simply agrees, adds unrelated noise, or resorts to isolated insults without addressing the argument). Because the text has already undergone strict normalization (noise removal and lowercasing), AI practitioners can directly feed this CSV into tokenizers and neural networks without needing to build complex data-cleaning pipelines from scratch.

本数据集(PolCSBD)旨在填补自然语言处理领域的一项关键空白:低资源混合代码语言中的政治反言论检测任务。我们的核心假设为:无法仅通过孤立的单条评论对反言论进行精准分类,其本质上需要结合前置发言的上下文语境。此外我们还假设,孟加拉国社交媒体用户在使用本土孟加拉语书写体系的同时,高频使用"孟加拉式英语(Banglish)"——一种以英语字母表拼音化孟加拉语词汇的语音混合语言,这为标准文本分类模型带来了显著障碍。 本数据集包含从政治讨论场景中提取的逾万条上下文对话文本对。每条数据均以直接对话结构组织,包含"父文本(parent_text)"与"回复文本(reply_text)"两个部分。该数据集展现了该地区复杂的语言生态:涵盖本土孟加拉语书写体系、全罗马化孟加拉语以及混合句式文本。其有效捕捉了互联网用户如何通过引用历史典故、使用攻击性辩论策略与讽刺手法来质疑政治叙事的行为模式。 数据集使用与解读说明: 本数据集经过深度优化,以机器学习就绪格式提供,非常适合用于训练、微调或基准测试Transformer模型(如mBERT、XLM-RoBERTa及BanglaBERT)的研究人员。 数据集仅包含三列: 1. "父文本(parent_text)":作为上下文基准的初始发言,已经过预处理以去除噪声。 2. "回复文本(reply_text)":对应回复发言,同样经过预处理。 3. "标签(label)":二元整数分类标签。标签值为'1'时代表反言论(Counter-Speech),即回复通过反驳、纠正或提出对立叙事来主动质疑父文本;标签值为'0'时代表非反言论(Non-Counter Speech),即回复仅表示认同、添加无关噪声,或仅进行孤立辱骂而未针对核心论点展开回应。 由于所有文本均已完成严格的归一化处理(含噪声去除与小写转换),AI从业者可直接将该CSV文件输入至分词器与神经网络中,无需从零构建复杂的数据清洗流程。

创建时间:
2026-04-27
二维码
社区交流群
二维码
科研交流群
商业服务