遇见数据集

StateParl

收藏
GESIS Data Catalogue2026-06-23 更新2026-06-25 收录
数据链接:
官方服务:

资源简介:

StateParl comprises the parliamentary speeches of members of parliament and government representatives in all 16 German state parliaments. The dataset consists of 16,168,903 paragraphs of parliamentary speech drawn from stenographic protocols covering plenary sessions from 1 January 2000 to 31 December 2025. StateParl integrates these protocols in an accessible, coherent, and machine-readable format that enables systematic applications across a wide range of academic disciplines. v3 introduces two new tables: a speeches table that groups consecutive paragraphs by speaker into addressable units and links non-speech content to the speeches they interrupt, and a mandates table that replaces the earlier identifier-mapping approach with a first-class entity storing resolved speaker identities across all 16 state parliaments. Speaker transition detection now combines state-specific regular expressions with a per-state machine-learning classifier, whose disagreements with the pattern-based approach drive iterative quality improvement. This documentation explains how StateParl was constructed, presents its final data structure, and provides the codebook. We also document the validation procedure, which confirmed 99.1% of sampled paragraphs as correct, and describe known limitations of the current release. StateParl comprises the parliamentary speeches of members of parliament and government representatives in all 16 German state parliaments. The dataset consists of 16,168,903 paragraphs of parliamentary speech drawn from stenographic protocols covering plenary sessions from 1 January 2000 to 31 December 2025. StateParl integrates these protocols in an accessible, coherent, and machine-readable format that enables systematic applications across a wide range of academic disciplines. v3 introduces two new tables: a speeches table that groups consecutive paragraphs by speaker into addressable units and links non-speech content to the speeches they interrupt, and a mandates table that replaces the earlier identifier-mapping approach with a first-class entity storing resolved speaker identities across all 16 state parliaments. Speaker transition detection now combines state-specific regular expressions with a per-state machine-learning classifier, whose disagreements with the pattern-based approach drive iterative quality improvement. This documentation explains how StateParl was constructed, presents its final data structure, and provides the codebook. We also document the validation procedure, which confirmed 99.1% of sampled paragraphs as correct, and describe known limitations of the current release.

提供机构:
GESIS Data Archive
创建时间:
2026-06-23
二维码
社区交流群
二维码
科研交流群
商业服务