Four Shades of Life Sciences (FSoLS)
收藏资源简介:
FSoLS数据集是一个新颖的、标记化的语料库,包含2,603篇关于14个生命科学主题的文章,从17个不同来源中检索,并分为四个生命科学出版物类别。数据集的设计旨在帮助机器学习模型识别和区分虚假信息文本。该数据集不仅包含完整的文章,而且涵盖了科学文本、通俗文本、替代科学文本和虚假信息文本等多种文本类型,从而为下游任务中的语言风格和内容分析提供了可能。FSoLS数据集的创建过程强调了平衡性,包括平衡的主题、数据来源和类别,以确保模型学习的是文本风格而非特定内容。该数据集的应用领域主要在于帮助用户在信息时代有效导航,特别是在健康和生命科学领域,识别和防止虚假信息的传播。
The FSoLS dataset is a novel, tokenized corpus consisting of 2,603 articles covering 14 life science topics, retrieved from 17 distinct sources and categorized into four life science publication categories. The dataset is designed to assist machine learning models in identifying and distinguishing disinformation texts. In addition to full-length articles, the dataset encompasses multiple text types including scientific texts, popular texts, alternative scientific texts, and disinformation texts, enabling linguistic style and content analysis for downstream tasks. The development of the FSoLS dataset emphasizes balance across topics, data sources, and categories, ensuring that models learn text styles rather than specific content. The primary application of this dataset is to help users effectively navigate the information age, especially in the health and life science domains, by identifying and preventing the spread of disinformation.

- 1Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life SciencesZB MED – Information Centre for Life Sciences · 2025年



