C-QUERI: Congressional Questions, Exchanges, and Responses in Institutions Dataset
收藏资源简介:
C-QUERI数据集是一个从美国第108届至第117届国会的委员会听证会记录中提取的问答交流数据集。该数据集包含从16130次听证会中提取的超过300万个语句,并对这些语句进行了分类和标注,为每个交流提供了全面的语料特征。该数据集旨在帮助学者研究国会议员如何寻求信息、构建议题以及如何对证人负责。数据集为研究国会行为、党派性以及民主提供了新的可能性,并为大型、自动化的问答分析提供了一个通用的框架。
The C-QUERI dataset is a question-and-answer (Q&A) conversational dataset extracted from the committee hearing transcripts of the 108th through 117th sessions of the United States Congress. This dataset contains over 3 million utterances extracted from 16,130 hearings, with these utterances categorized and annotated to provide comprehensive corpus features for each conversational exchange. It aims to assist scholars in studying how members of Congress seek information, frame policy issues, and hold witnesses accountable. The dataset offers new opportunities for research on congressional behavior, partisanship, and democracy, as well as a general framework for large-scale, automated Q&A analysis.




