遇见数据集

Training data for the shared task Ideology and Power Identification in Parliamentary Debates (2024)

收藏
Zenodo2025-01-04 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains a selection of speeches from ParlaMint corpora (version 4.0) as the training set for the shared task on "Ideology and Power Identification in Parliamentary Debates" in CLEF 2024. All files are tab-separated text files with the following fields: "id" is a unique (arbitrary) ID for each text. "speaker" is a unique (arbitrary) ID for each speaker. There may be multiple speeches from the same speaker. "sex" is the (binary/biological) sex of the speaker. This information is collected from varying sources (typically data published by the respective parliament), and in some cases it may be unspecified or unknown. "text" is the transcribed text of the parliamentary speech. Real examples may include line breaks, and other special sequences escaped or quoted. "text_en" is an automatic English translation of the corresponding text. This field may be empty (obviously) for speeches in English, but the translations may be missing for a small number of non-English speeches as well. "label" is the binary/numeric label. For political orientation, 0 is left and 1 is right. For power identification 0 indicates coalition (or governing party) and 1 indicates opposition. File names indicate the task and the parliament. We provide data from the following national and regional parliaments. Austria (at) Bosnia and Herzegovina (ba) Belgium (be) Bulgaria (bg) Czechia (cz) Denmark (dk) Estonia (ee) [only political orientation] Spain (es) Catalonia (es-ct) Galicia (es-ga) Basque Country (es-pv) [only power] Finland (fi) France (fr) Great Britain (gb) Greece (gr) Croatia (hr) Hungary (hu) Iceland (is) [only political orientation] Italy (it) Latvia (lv) The Netherlands (nl) Norway (no) [only political orientation] Poland (pl) Portugal (pt) Serbia (rs) Sweden (se) [only political orientation] Slovenia (si) Turkey (tr) Ukraine (ua) The number of training instances and the class imbalance differs for each training set. We do not provide a fixed validation split. Please see the shared task website for further description of the data set and the sampling process.

本数据集选取ParlaMint语料库(版本4.0)中的多篇演讲作为训练集,用于CLEF 2024「议会辩论中的意识形态与权力识别」共享任务。所有文件均为制表符分隔的文本文件,包含以下字段: "id" 为每篇文本的唯一(任意)标识符。 "speaker" 为每位发言人的唯一(任意)标识符,同一发言人可发表多篇演讲。 "sex" 表示发言人的(二元/生理)性别。此类信息源自多种渠道(通常为各议会公开数据),部分情况下可能未指定或未知。 "text" 为议会演讲的转写文本。实际示例中可能包含换行符,其余特殊序列均会经过转义或引用处理。 "text_en" 为对应文本的自动英文翻译。英语演讲的该字段显然可为空,少量非英语演讲也可能缺失翻译。 "label" 为二元/数值标签。若用于政治倾向分类,0代表左翼,1代表右翼;若用于权力识别任务,0表示联合(或执政)党派,1表示反对党。 文件名标注了任务与所属议会。我们提供以下国家及地区议会的数据集: - 奥地利(at) - 波斯尼亚和黑塞哥维那(ba) - 比利时(be) - 保加利亚(bg) - 捷克(cz) - 丹麦(dk) - 爱沙尼亚(ee)[仅提供政治倾向分类数据] - 西班牙(es) - 加泰罗尼亚(es-ct) - 加利西亚(es-ga) - 巴斯克地区(es-pv)[仅提供权力识别数据] - 芬兰(fi) - 法国(fr) - 英国(gb) - 希腊(gr) - 克罗地亚(hr) - 匈牙利(hu) - 冰岛(is)[仅提供政治倾向分类数据] - 意大利(it) - 拉脱维亚(lv) - 荷兰(nl) - 挪威(no)[仅提供政治倾向分类数据] - 波兰(pl) - 葡萄牙(pt) - 塞尔维亚(rs) - 瑞典(se)[仅提供政治倾向分类数据] - 斯洛文尼亚(si) - 土耳其(tr) - 乌克兰(ua) 各训练集的样本数量与类别不平衡程度各不相同,我们未提供固定的验证集划分方式。如需了解数据集与采样流程的更多细节,请访问共享任务官方网站。

提供机构:
Zenodo
创建时间:
2024-01-02
二维码
社区交流群
二维码
科研交流群
商业服务