TBMM Parliamentary Proceedings Corpus with Speaker-Turn Segmentation, Party Linkage, and Validated Metadata, 1950-2023
收藏资源简介:
A speech-level corpus of Turkish Grand National Assembly (TBMM) parliamentary proceedings covering 4 January 1950 through 23 April 2023 — the full period of competitive multi-party parliamentary politics in Turkey. The corpus contains 2,418,313 speaker turns across five chambers (TBMM, Millet Meclisi, Senato, Danışma Meclisi, MGK), of which 702,503 are linked to specific members of parliament and parties via a Wikipedia-derived roster of 10,740 MP-term records with gender annotations. An analytical core subset of 136,118 high-confidence floor speeches (49.4 million words) is provided for quantitative political analysis. This subset covers 2,868 unique MPs across 26 political parties and is filtered to include only substantive speeches with exact name+province+decade linkage at HIGH or MEDIUM confidence. The corpus combines two sources. For 1950-2018, it extends the Turkronicles OCR archive (Yazar et al. 2025, CC BY 4.0) with additional speaker-turn segmentation and linkage layers. For the 27th Legislative Term (2018-2023), it adds 314,899 speaker turns extracted directly from 537 official Word-format session transcripts published by TBMM. Speaker boundaries are detected using a ten-pattern unified parser with decade-aware OCR normalization. A targeted correction in this release recovers more than 6,000 multi-word party-group speeches that earlier pipelines misclassified as individual MP turns. Note: This Harvard Dataverse deposit is a mirror of the primary Zenodo release (https://doi.org/10.5281/zenodo.19713325). Please use the Zenodo DOI as the primary citation reference.
本语料库为土耳其大国民议会(Turkish Grand National Assembly,TBMM)议事话语级语料库,涵盖1950年1月4日至2023年4月23日的土耳其全部竞争性多党议会政治时期。该语料库涵盖5个议事机构下的2,418,313条发言轮次,分别为TBMM、Millet Meclisi、Senato、Danışma Meclisi、MGK;其中702,503条发言轮次可通过维基百科衍生的10,740条带性别标注的议员任期记录名册,关联至特定议员与政党。为支撑量化政治分析,本语料库附带由136,118条高置信度正式议事发言构成的分析核心子集,总词量达4940万。该子集涵盖隶属于26个政党的2,868名不同议员,且仅筛选出置信度为高或中等、具备精确姓名-省份-年代关联的实质性发言。本语料库整合了两类数据源:针对1950年至2018年的内容,本语料库在Turkronicles光学字符识别(Optical Character Recognition,OCR)档案(Yazar等,2025,CC BY 4.0协议)基础上,新增了发言轮次切分与关联层级;针对第27届立法任期(2018年-2023年),本语料库新增了314,899条发言轮次,这些轮次直接提取自TBMM发布的537份官方Word格式议事实录。发言边界的识别采用了十模式统一解析器,并结合了年代感知的OCR归一化处理。本次发布还包含针对性修正内容,修复了6,000余条此前被错误分类为议员个人发言轮次的多词政党团组发言。注:本哈佛大学数据视界(Harvard Dataverse)存档为Zenodo官方发布版本的镜像(https://doi.org/10.5281/zenodo.19713325),请以Zenodo的DOI作为主要引用标识。



