遇见数据集

Rhetoric of Division: a labeled corpus of 5.6 million U.S. Congressional speeches and tweets, 2008 to 2025

收藏
Zenodo2026-04-08 更新2026-05-26 收录
官方服务:

资源简介:

A research-grade dataset of 645,524 U.S. Congressional floor speeches and 4,972,551 tweets from sitting members of Congress, spanning January 2008 through October 2025. Every record is labeled by a four-model Large Language Model ensemble (Gemini 2.5 Flash Lite, GPT-5 Nano, Grok 4.1 Fast, and a fine-tuned ModernBERT) across five domains of polarizing rhetoric: group identity and boundaries, conflict framing and affective polarization, delegitimization and dehumanization, anti-deliberation and absolutism, and system threats. Each record carries the original cleaned text, the speaker's Bioguide identifier for joining with member metadata, the per-model classifications, and an ensemble vote. The corpus is the largest open dataset of LLM-classified U.S. political communication that we are aware of and was assembled to support the Master's thesis "Quantifying Political Polarization Using Text Mining" (Nova School of Business and Economics, January 2026, final grade 19/20). Data is organized as one JSONL file per platform per month. Member metadata is included as a CSV. The full schema, methodology, and provenance are documented in the README inside the archive and in the thesis PDF, which lives in the companion GitHub repository: https://github.com/FelixSBuehrm/rhetoric-of-division

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务