MARIANNE-INRIA/PoliticalDebatesCorpus-EN
收藏资源简介:
PoliticalDebatesCorpus-EN(也称为RooseDebates)是一个大规模英语政治辩论和演讲转录本语料库,专为领域特定语言模型预训练而构建。它是专为政治辩论的对话和论证性质定制的语言模型RooseBERT的训练语料库。该语料库汇集了来自多个国家和国际机构的电视总统辩论和议会辩论,总计约11 GB的文本,时间跨度为1946年至2025年。数据来源包括非洲议会辩论(加纳和南非)、澳大利亚议会辩论、加拿大议会辩论、欧洲议会辩论、爱尔兰议会辩论、新西兰议会辩论、苏格兰议会辩论、英国议会辩论、联合国大会辩论、联合国安理会辩论和美国辩论等权威政治设置。语料库经过预处理,移除了超链接和标记标签,并合并了多个空格。预期用途包括政治话语分析的语言模型预训练和微调,支持下游任务如立场检测、情感分析、论证挖掘和政策分类。但语料库仅限英语,地理和语言覆盖不均,且可能包含偏见或攻击性语言。
PoliticalDebatesCorpus-EN is a large-scale English corpus of political debate and speech transcripts, assembled for domain-specific language model pre-training. It is the training corpus behind RooseBERT, a language model tailored to the dialogic and argumentative nature of political debates. The corpus brings together televised presidential debates and parliamentary debates from multiple countries and international bodies, totalling ~11 GB of text and spanning the period 1946–2025. Data sources include African Parliamentary Debates (Ghana & South Africa), Australian Parliamentary Debates, Canadian Parliamentary Debates, European Parliamentary Debates, Irish Parliamentary Debates, New Zealand Parliamentary Debates, Scottish Parliamentary Debates, UK Parliamentary Debates, UN General Debate Corpus, UN Security Council Debates, and United States Debates, all from authoritative political settings. The corpus was pre-processed to remove hyperlinks and markup tags and to collapse multiple spaces. Intended uses include pre-training and fine-tuning of language models for political discourse analysis, supporting downstream tasks such as stance detection, sentiment analysis, argument mining, and policy classification. However, the corpus is English-only with uneven geographic and linguistic coverage, and may contain biased or offensive language.




