遇见数据集

AURORA-MESS: Annotated Uncertainty and Reference in Open Research Articles for Multidisciplinary and Empirical Social Science

收藏
Zenodo2025-03-11 更新2026-05-26 收录
官方服务:

资源简介:

The AURORA-MESS dataset offers valuable insights into the representation of uncertainty in scientific literature across various domains, along with associated authorial references. Researchers and practitioners can use this dataset to study and analyze the variations of uncertainty expressions in scholarly discourse. This dataset contains sentences extracted from articles in a wide range of fields, covering both Science, Technology, and Medicine (STM); and Social Sciences and Humanities (SSH) and annotated with respect to uncertainty in science. The dataset is derived from PubMed, Scopus, Web of Science (WoS) and Social Science Open Access Repository (SSOAR). It has been produced as part of the ANR InSciM (Modelling Uncertainty in Science) project. For a more comprehensive understanding of the construction of the dataset, including the selection of journals, sampling procedure, and the annotation methodology, see Ningrum and Atanassova (2024); and Ningrum, Mayr, and Atanassova (2023). The dataset is provided in CSV format. The columns in the table are as follows: sentence_id: A unique internal identifier for each sentence. journal_name: The name of the journal in which the article was published. article_title: The title of the article from which the sentence was extracted. publication_year: The year the article was published. document_id: The URL where the article is published. sentence: The text of the sentence. uncertainty: 'Y' if the sentence expresses uncertainty, and 'N' otherwise. reference: This column contains three classes: '1' indicates uncertainty arising directly from the authors’ statements. '2' represents uncertainty attributed to prior research, where the authors cite past studies to introduce uncertainty. '3' includes instances where the author(s) express shared uncertainty, drawing on both their findings or arguments and previous research. Reference Ningrum, P. K., & Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. Scientometrics, 1-25. https://doi.org/10.1007/s11192-024-05009-z Ningrum, P. K., Mayr, P., & Atanassova, I. (2023). UnScientify: Detecting Scientific Uncertainty in Scholarly Full Text. arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2307.14236 Acknowledgment We would like to acknowledge GESIS - Leibniz Institute for the Social Sciences for providing access to SSOAR, which has been instrumental in supporting the creation of the dataset.

AURORA-MESS数据集为剖析多领域科学文献中的不确定性表征及其相关作者引用提供了极具价值的研究视角。研究人员与从业者可借助该数据集,探究并分析学术话语中不确定性表达的变体形式。 本数据集提取自涵盖科技医学(Science, Technology, and Medicine, STM)以及人文社会科学(Social Sciences and Humanities, SSH)在内的众多领域的学术文章,并针对科学领域的不确定性进行了标注。该数据集源自PubMed、Scopus、Web of Science(WoS)以及社会科学开放获取库(Social Science Open Access Repository, SSOAR),是ANR InSciM(*Modelling Uncertainty in Science*,科学不确定性建模)项目的产出成果之一。若需深入了解该数据集的构建细节,包括期刊遴选、采样流程与标注方法,请参阅Ningrum与Atanassova(2024)以及Ningrum、Mayr与Atanassova(2023)的相关研究。 本数据集以CSV格式提供。数据集表格包含以下字段: sentence_id:每一条句子的唯一内部标识符。 journal_name:文章发表所在的期刊名称。 article_title:提取该句子的文章标题。 publication_year:文章的发表年份。 document_id:文章发表的URL链接。 sentence:句子原文内容。 uncertainty:若句子表达不确定性则标注为'Y',否则标注为'N'。 reference:该字段包含三类标注: '1' 表示不确定性直接源自作者本人的陈述; '2' 表示不确定性源于既往研究,即作者通过引用过往研究来引入不确定性; '3' 涵盖作者结合自身研究发现/论证与既往研究,表达共同不确定性的场景。 参考文献 Ningrum, P. K. & Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. *Scientometrics*, 1-25. https://doi.org/10.1007/s11192-024-05009-z Ningrum, P. K., Mayr, P. & Atanassova, I. (2023). UnScientify: Detecting Scientific Uncertainty in Scholarly Full Text. arXiv [Cs.CL]. 取自 http://arxiv.org/abs/2307.14236 致谢 衷心感谢莱布尼茨社会科学研究所GESIS为我们提供SSOAR的访问权限,该资源对本数据集的构建发挥了至关重要的支撑作用。

提供机构:
Zenodo
创建时间:
2025-03-11
二维码
社区交流群
二维码
科研交流群
商业服务