遇见数据集

runetrust/blame-folketinget-dk

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为Folketinget Denmark with Blame Labels,是一个丹麦政治辩论文本数据集,专注于指责标签分类。数据集包含丹麦议会(Folketinget)的演讲文本,时间跨度从1997年10月7日至2018年12月20日(基于ParlSpeech V2数据集)以及2019年1月8日至2026年2月26日(从Folketinget的sFTP服务器获取)。数据经过机器翻译(使用opus-MT)后,通过Political DEBATE模型生成了指责标签(Blame: 1, No Blame: 0)。数据集特征包括文本、标签、演讲者、日期、政党、段落编号和句子编号。分割包括训练集(450,000个示例)、验证集(50,000个示例)、测试集(424个示例,手动标注,标注者间一致性为84.8%,Cohens Kappa为0.675)和推理集(5,098,494个示例,为清理后剩余数据)。数据集大小为约1.19 GB,下载大小为464 MB,许可证为CC0 1.0,适用于政治和辩论相关研究。

The dataset Folketinget Denmark with Blame Labels is a Danish political debate text dataset focused on blame label classification. It contains speech texts from the Danish Parliament (Folketinget), spanning from October 7, 1997, to December 20, 2018 (based on the ParlSpeech V2 dataset) and from January 8, 2019, to February 26, 2026 (fetched from Folketingets sFTP server). After machine translation (using opus-MT), blame labels (Blame: 1, No Blame: 0) were generated using the Political DEBATE model. Features include text, labels, speaker, date, party, paragraph number, and sentence number. Splits consist of a training set (450,000 examples), validation set (50,000 examples), test set (424 examples, manually annotated with inter-annotator agreement of 84.8%, Cohens Kappa of 0.675), and inference set (5,098,494 examples, representing remaining data after cleaning). The dataset size is approximately 1.19 GB, with a download size of 464 MB, licensed under CC0 1.0, and suitable for politics and debates research.

提供机构:
runetrust
二维码
社区交流群
二维码
科研交流群
商业服务