遇见数据集

SKILL-IR-Discourse

收藏
Zenodo2026-03-05 更新2026-05-26 收录
官方服务:

资源简介:

We present a large annotated corpus of scholarly discourse in the domain of International Relations, a subfield of political science. The corpus comprises 190 articles (over 1500K tokens) annotated at the argumentation, basic rhetorical, and domain level. Five of the included articles (ca. 62K tokens) constitute a Gold-standard, coded by domain experts. The remaining articles were coded by annotators trained on the Gold-standard and monitored for annotation quality. We describe our corpus creation methodology, the annotation process and quality assurance, the corpus itself, and present insights into the data: Most argumentative structures in the data are simple premise-conclusion structures, fewer than half of the claims have explicit supporting evidence. Counter-arguments to claims are rare. The claim-to-support ratio varies widely between articles; possibly to some extent due to the topics covered (with clear common ground) or to the differences between authors' styles. The distribution of theoretical vs. evaluative statements varies strongly between articles; this can be attributed to such factors as different methodological approaches between the articles and the methodological focus of the publishing journal.

本研究构建了一个大型政治学下属子领域国际关系领域的学术话语标注语料库(corpus)。该语料库包含190篇学术文章,总词元(Token)数超过150万,在论证、基础修辞以及领域三个层面完成标注。其中5篇文章(约6.2万词元)构成金标准(Gold-standard)数据集,由领域专家进行编码标注;剩余文章则由经过金标准数据集培训的标注人员完成编码,并对标注质量进行全程监控。本研究详细阐述了该语料库的构建方法、标注流程与质量保障机制,同时对数据集本身展开分析,并呈现如下数据洞察:数据集中绝大多数论证结构为简单的前提-结论结构,仅有不到半数的主张带有明确的支撑性证据;针对主张的反论极为少见。不同文章间的主张-支撑比差异显著,这在一定程度上可能源于所涉主题(存在明确的共识基础)的差异,或是作者写作风格的不同。不同文章中理论性陈述与评价性陈述的分布差异极大,这可归因于文章间研究方法的差异,以及发表期刊的方法论侧重等因素。

提供机构:
Zenodo
创建时间:
2025-10-17
二维码
社区交流群
二维码
科研交流群
商业服务