SpeakGer
收藏资源简介:
SpeakGer数据集是由多特蒙德工业大学统计系创建的一个丰富的元数据语音语料库,涵盖了德国16个联邦州议会和德国联邦议院从1947年到2023年的辩论记录。该数据集包含10,806,105条演讲,每条演讲都附有丰富的元数据,如演讲者的党派、年龄、选区及其党派的政治立场,以及听众对演讲的反应等信息。数据集的创建过程包括从各议会网站收集数据,并使用OCR技术提取和校正文本。SpeakGer数据集旨在通过提供详细的元数据,支持政治科学领域的精细化研究,特别是在自然语言处理和政治文本分析方面。
SpeakGer Dataset is a metadata-rich speech corpus developed by the Department of Statistics, TU Dortmund University. It covers debate transcripts from the 16 state parliaments of Germany and the German Bundestag spanning from 1947 to 2023. The dataset contains 10,806,105 speeches, each paired with comprehensive metadata including the speaker's political party, age, constituency, the political stance of their affiliated party, and audience reactions to the speech, among other relevant details. The dataset was constructed by collecting raw data from official parliamentary websites and employing OCR technology to extract and correct the textual content. The SpeakGer Dataset aims to support fine-grained research in political science, especially within the domains of natural language processing and political text analysis, by providing detailed and comprehensive metadata.

- 1SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments多特蒙德工业大学统计系 · 2024年



