siddqamar/GMO-Myths-and-Truths
收藏资源简介:
该数据集包含关于基因改造生物(GMOs)的声明和基于证据的发现的结构化集合,数据来源于技术报告《GMO Myths and Truths: An evidence-based examination of the claims made for the safety and efficacy of genetically modified crops》(版本1.3a,2012年6月)。数据集设计用于生物技术领域的二元文本分类、情感分析和语义搜索任务。数据集分为两个主要子集:1. 纯提取:直接从源文档中提取的神话/真相配对声明(94个平衡样本);2. 增强版本:通过语言数据增强(包括同义词替换、结构变异和上下文包装)扩展的鲁棒训练集(500多个平衡样本)。标签分类为二元格式:0代表神话(行业支持者的声明或营销论点),1代表真相(基于证据的发现、科学反驳或安全数据)。数据集支持的任务包括二元文本分类、语义相似性和数据增强研究。数据集是基于Earth Open Source出版物的衍生作品,仅供研究、基准测试和非商业教育用途。
This dataset contains a structured collection of claims and evidence-based findings regarding Genetically Modified Organisms (GMOs). The data was extracted and adapted from the technical report: "GMO Myths and Truths: An evidence-based examination of the claims made for the safety and efficacy of genetically modified crops" (Version 1.3a, June 2012). It is designed for binary text classification, sentiment analysis, and semantic search tasks within the biotechnology domain. The dataset is organized into two primary subsets: 1. Pure: Direct extractions of paired Myth/Truth statements from the source document (94 balanced samples). 2. Augmented: A robust training set expanded via Linguistic Data Augmentation (including synonym substitution, structural variation, and contextual wrapping) to improve model generalization (500+ balanced samples). Statements are categorized into a binary format: 0 represents Myth (industry-proponent claims or marketing arguments), and 1 represents Truth (evidence-based findings, scientific rebuttals, or safety data). Supported tasks include Binary Text Classification, Semantic Similarity, and Data Augmentation Research. The dataset is a derivative work based on the Earth Open Source publication and is provided for research, benchmarking, and non-commercial educational use.




