遇见数据集

Annotated Dataset for Uncertainty Mining : Gold Standard

收藏
Zenodo2024-11-13 更新2026-05-26 收录
官方服务:

资源简介:

Description of the dataset In order to study the expression of uncertainty in scientific articles, we have put together an interdisciplinary corpus of journals in the fields of Science, Technology and Medicine (STM) and the Humanities and Social Sciences (SHS). The selection of journals in our corpus is based on the Scimago Journal and Country Rank (SJR) classification, which is based on Scopus, the largest academic database available online. We have selected journals covering various disciplines, such as medicine, biochemistry, genetics and molecular biology, computer science, social sciences, environmental sciences, psychology, arts and humanities. For each discipline, we selected the five highest-ranked journals. In addition, we have included the journals PLoS ONE and Nature, both of which are interdisciplinary and highly ranked. Based on the corpus of articles from different disciplines described above, we created a set of annotated sentences as follows: 593 were pre-selected automatically, by studying the occurrences of the lists of uncertainty indices proposed by Bongelli et al. (2019), Chen et al. (2018) and Hyland (1996). The remaining sentences were extracted from a subset of articles, consisting of two randomly selected articles per journal. These articles were examined by two human annotators to identify sentences containing uncertainty and to annotate them. 600 sentences not expressing scientific uncertainty were manually identified and reviewed by two annotators The sentences were annotated by two independent annotators following the annotation guide proposed by Ningrum and Atanassova (2024). The annotators were trained on the basis of an annotation guide and previously annotated sentences in order to guarantee the consistency of the annotations. Each sentence was annotated as expressing or not expressing uncertainty (Uncertainty and No Uncertainty).Sentences expressing uncertainty were then annotated along five dimensions: Reference , Nature, Context , Timeline and Expression. The annotators reached an average agreement score of 0.414 according to Cohen's Kappa test, which shows the difficulty of the task of annotating scientific uncertainty.Finally, conflicting annotations were resolved by a third independent annotator. Our final corpus thus consists of a total of 1 840 sentences from 496 articles in 21 English-language journals from 8 different disciplines.The columns of the table are as follows: journal: name of the journal from where the article originates article_title: title of the article from where the sentence is extracted publication_year: year of publication of the article sentence_text: text of the sentence expressing or not expressing uncertainty uncertainty: 1 if the sentence expresses uncertainty and 0 otherwise; ref, nature, context, timeline, expression: annotations of the type of uncertainty according to the annotation framework proposed by Ningrum and Atanassova (2023). The annotation of each dimension in this dataset are in numeric format rather than textual. The mapping betwen textual and numeric labels is presented in the Table below. Dimension 1 2 3 4 5 Reference Author Former Both Nature Epistemic Aleatory Both Context Background Methods Res&Disc Conclusion Others Timeline Past Present Future Expression Quantified Unquantified This gold standard has been produced as part of the ANR InSciM (Modelling Uncertainty in Science) project. References Bongelli, R., Riccioni, I., Burro, R., & Zuczkowski, A. (2019). Writers’ uncertainty in scientific and popular biomedical articles. A comparative analysis of the British Medical Journal and Discover Magazine [Publisher: Public Library of Science]. PLoS ONE, 14 (9). https://doi.org/10.1371/journal.pone.0221933 Chen, C., Song, M., & Heo, G. E. (2018). A scalable and adaptive method for finding semantically equivalent cue words of uncertainty. Journal of Informetrics, 12 (1), 158–180. https://doi.org/10.1016/j.joi.2017.12.004 Hyland, K. E. (1996). Talking to the academy forms of hedging in science research articles [Publisher: SAGE Publications Inc.]. Written Communication, 13 (2), 251–281. https://doi.org/10.1177/0741088396013002004 Ningrum, P. K., & Atanassova, I. (2023). Scientific Uncertainty: An Annotation Framework and Corpus Study in Different Disciplines. 19th International Conference of the International Society for Scientometrics and Informetrics (ISSI 2023). https://doi.org/10.5281/zenodo.8306035 Ningrum, P. K., & Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. Scientometrics. https://doi.org/10.1007/s11192-024-05009-z

数据集描述 为研究科学论文中不确定性的表达形式,我们构建了涵盖科技医学(Science, Technology and Medicine, STM)与人文社会科学(Humanities and Social Sciences, SHS)领域期刊的跨学科语料库。本语料库的期刊遴选基于Scimago期刊与国家排名(Scimago Journal and Country Rank, SJR)分类体系,该体系依托当前全球最大的在线学术数据库Scopus搭建。我们选取了覆盖多学科的期刊,包括医学、生物化学、遗传学与分子生物学、计算机科学、社会科学、环境科学、心理学、艺术与人文领域。针对每个学科,我们遴选了排名前五的期刊。此外,我们还纳入了兼具跨学科特性与高影响力的期刊《公共科学图书馆·综合》(PLoS ONE)与《自然》(Nature)。 基于上述跨学科文章语料库,我们构建了带标注的句子集,流程如下: 首先,通过检索Bongelli等人(2019)、Chen等人(2018)与Hyland(1996)提出的不确定性标识词列表,自动预筛选出593个句子。 剩余句子则从期刊子集内提取,该子集包含每本期刊随机选取的2篇文章。由两名人工标注员对这些文章进行审阅,识别出包含不确定性表达的句子并完成标注。 另有600个未表达科学不确定性的句子,由两名标注员手动甄别并审核。 所有句子均由两名独立标注员依据Ningrum与Atanassova(2024)提出的标注指南完成标注。标注员先通过学习标注指南与已标注样本开展培训,以保障标注一致性。每句被标注为「表达不确定性」或「未表达不确定性」两类。针对表达不确定性的句子,进一步从五个维度进行标注:参照维度(Reference)、性质维度(Nature)、语境维度(Context)、时间维度(Timeline)与表达形式维度(Expression)。根据科恩卡帕(Cohen's Kappa)检验,标注员间的平均一致性得分为0.414,这一结果反映了科学不确定性标注任务的难度。最终,由第三名独立标注员对存在标注冲突的样本进行仲裁。 因此,本最终语料库共包含来自8个不同学科、21种英文期刊的496篇文章中的1840个句子。数据集表格的字段说明如下: - journal:文章来源期刊名称 - article_title:句子提取自的文章标题 - publication_year:文章发表年份 - sentence_text:表达或未表达不确定性的句子文本 - uncertainty:若句子表达不确定性则取值为1,否则为0 - ref、nature、context、timeline、expression:依据Ningrum与Atanassova(2023)提出的标注框架所标注的不确定性类型。本数据集内各维度的标注均采用数值格式而非文本格式,文本标签与数值标签的对应关系如下表所示: | 维度 | 1 | 2 | 3 | 4 | 5 | |--------------------|------------|----------|--------|------------|--------| | Reference(参照) | 作者(Author) | 前者(Former) | 两者(Both) | / | / | | Nature(性质) | 认知型(Epistemic) | 随机型(Aleatory) | 两者(Both) | / | / | | Context(语境) | 背景(Background) | 方法(Methods) | 研究与讨论(Res&Disc) | 结论(Conclusion) | 其他(Others) | | Timeline(时间) | 过去(Past) | 现在(Present) | 未来(Future) | / | / | | Expression(表达) | 量化型(Quantified) | 非量化型(Unquantified) | / | / | / | 本金标准数据集(gold standard)作为ANR InSciM项目(科学不确定性建模,Modelling Uncertainty in Science)的产出成果完成构建。 参考文献 1. Bongelli, R., Riccioni, I., Burro, R., & Zuczkowski, A. (2019). 科研与科普生物医学文章中的作者不确定性:《英国医学期刊》与《发现》杂志的比较分析 [出版商:公共科学图书馆]. PLoS ONE, 14(9). https://doi.org/10.1371/journal.pone.0221933 2. Chen, C., Song, M., & Heo, G. E. (2018). 一种可扩展且自适应的语义等价不确定性提示词挖掘方法. 信息计量学期刊, 12(1), 158–180. https://doi.org/10.1016/j.joi.2017.12.004 3. Hyland, K. E. (1996). 与学术共同体对话:科研论文中的模糊限定表达 [出版商:SAGE出版集团]. Written Communication, 13(2), 251–281. https://doi.org/10.1177/0741088396013002004 4. Ningrum, P. K., & Atanassova, I. (2023). 科学不确定性:跨学科标注框架与语料库研究. 第19届国际科学计量学与信息计量学学会国际会议(ISSI 2023). https://doi.org/10.5281/zenodo.8306035 5. Ningrum, P. K., & Atanassova, I. (2024). 基于语言模式的科学不确定性标注. 科学计量学, https://doi.org/10.1007/s11192-024-05009-z

提供机构:
Zenodo
创建时间:
2024-11-13
二维码
社区交流群
二维码
科研交流群
商业服务