遇见数据集

bigbio/mirna

收藏
Hugging Face2022-12-22 更新2024-03-04 收录
官方服务:

资源简介:

该数据集包含301个Medline引用,这些文档的摘要文本中提到了miRNA。基因、疾病和miRNA实体被手动标注。数据集分为训练集和测试集,分别来自201和100个文档。数据集的主要任务是命名实体识别(NER)和命名实体消歧(NED)。数据集是单语言的,仅包含英文内容,并且是公开的,可以在PubMed上找到。数据集的许可证是CC BY-NC 3.0。

This dataset comprises 301 Medline citations, with miRNAs mentioned in the abstracts of all included documents. Entities of genes, diseases, and miRNAs were manually annotated. The dataset is split into training and test sets, which are sourced from 201 and 100 documents respectively. The core tasks of this dataset are Named Entity Recognition (NER) and Named Entity Disambiguation (NED). This is a monolingual dataset containing only English content, and it is publicly available and retrievable on PubMed. The dataset is licensed under CC BY-NC 3.0.

提供机构:
bigbio
原始信息汇总

数据集概述

基本信息

  • 名称: miRNA
  • 语言: 英语
  • 许可证: CC BY NC 3.0
  • 多语言支持: 单语(英语)
  • PubMed可用性: 是
  • 公开可用性: 是

数据集内容

  • 文档数量: 301篇Medline文献
  • 文件组成: 分为训练集(201篇文献)和测试集(100篇文献)
  • 标注内容: 手动标注了基因、疾病和miRNA实体

任务类型

  • 命名实体识别 (NER)
  • 命名实体消歧 (NED)

引用信息

@Article{Bagewadi2014, author={Bagewadi, Shweta and Bobi{{c}}, Tamara and Hofmann-Apitius, Martin and Fluck, Juliane and Klinger, Roman}, title={Detecting miRNA Mentions and Relations in Biomedical Literature}, journal={F1000Research}, year={2014}, month={Aug}, day={28}, publisher={F1000Research}, volume={3}, pages={205-205}, keywords={MicroRNAs; corpus; prediction algorithms}, abstract={ INTRODUCTION: MicroRNAs (miRNAs) have demonstrated their potential as post-transcriptional gene expression regulators, participating in a wide spectrum of regulatory events such as apoptosis, differentiation, and stress response. Apart from the role of miRNAs in normal physiology, their dysregulation is implicated in a vast array of diseases. Dissection of miRNA-related associations are valuable for contemplating their mechanism in diseases, leading to the discovery of novel miRNAs for disease prognosis, diagnosis, and therapy. MOTIVATION: Apart from databases and prediction tools, miRNA-related information is largely available as unstructured text. Manual retrieval of these associations can be labor-intensive due to steadily growing number of publications. Additionally, most of the published miRNA entity recognition methods are keyword based, further subjected to manual inspection for retrieval of relations. Despite the fact that several databases host miRNA-associations derived from text, lower sensitivity and lack of published details for miRNA entity recognition and associated relations identification has motivated the need for developing comprehensive methods that are freely available for the scientific community. Additionally, the lack of a standard corpus for miRNA-relations has caused difficulty in evaluating the available systems. We propose methods to automatically extract mentions of miRNAs, species, genes/proteins, disease, and relations from scientific literature. Our generated corpora, along with dictionaries, and miRNA regular expression are freely available for academic purposes. To our knowledge, these resources are the most comprehensive developed so far. RESULTS: The identification of specific miRNA mentions reaches a recall of 0.94 and precision of 0.93. Extraction of miRNA-disease and miRNA-gene relations lead to an F1 score of up to 0.76. A comparison of the information extracted by our approach to the databases miR2Disease and miRSel for the extraction of Alzheimers disease related relations shows the capability of our proposed methods in identifying correct relations with improved sensitivity. The published resources and described methods can help the researchers for maximal retrieval of miRNA-relations and generation of miRNA-regulatory networks. AVAILABILITY: The training and test corpora, annotation guidelines, developed dictionaries, and supplementary files are available at http://www.scai.fraunhofer.de/mirna-corpora.html. }, note={26535109[pmid]}, note={PMC4602280[pmcid]}, issn={2046-1402}, url={https://pubmed.ncbi.nlm.nih.gov/26535109}, language={eng} }

搜集汇总
数据集介绍
bigbio/mirna 数据集图片
构建方式
在生物医学文本挖掘领域,非编码RNA尤其是微小RNA(miRNA)的实体识别与关系抽取是理解其调控机制的关键。该数据集基于301篇Medline文献摘要构建,由专业标注人员对其中涉及的miRNA、基因、疾病三类实体进行人工注释。语料库被划分为训练集(201篇)和测试集(100篇)两个独立文件,确保了模型评估的公正性与可重复性。构建过程中严格遵循标注指南,并辅以词典与正则表达式工具,形成了兼具规模与质量的标准资源。
特点
该数据集的核心特点在于其多任务兼容性,同时支持命名实体识别与实体消歧两项任务。涵盖的实体类型包括miRNA、基因/蛋白质及疾病,覆盖了生物医学文献中miRNA相关研究的主要语义单元。数据来源于经过PubMed索引的真实文献,具有高度的领域相关性与专业代表性。此外,语料库的公开性与CC-BY-NC-3.0许可协议,为学术研究提供了自由使用的便利,促进了miRNA信息抽取方法的可复现与比较。
使用方法
使用时,用户可直接加载训练集与测试集文件进行模型训练与评估,适用于基于序列标注的NER模型或基于图神经网络的NED系统。数据格式支持直接转换为BIO或BILOU标签体系,便于接入主流深度学习框架。推荐结合原始论文中提供的词典与正则表达式进行预训练或特征增强,以提升实体识别的召回率与精确率。同时,该语料库可作为miRNA-疾病与miRNA-基因关系抽取的基准资源,用于验证信息检索系统的性能。
背景与挑战
背景概述
微小RNA(miRNA)作为基因表达的关键转录后调控因子,在细胞凋亡、分化和应激反应等生物学过程中扮演着核心角色,其失调与多种疾病的发生发展密切相关。然而,随着生物医学文献数量的爆炸性增长,从非结构化文本中高效、准确地挖掘miRNA相关关联信息成为一项严峻挑战。为应对这一需求,Fraunhofer SCAI研究所的Shweta Bagewadi、Tamara Bobić、Martin Hofmann-Apitius、Juliane Fluck及Roman Klinger等研究人员于2014年创建了miRNA语料库。该数据集精选了301篇Medline文献摘要,由201篇训练文档和100篇测试文档构成,并由领域专家手动标注了基因、疾病及miRNA实体。这一开创性工作为miRNA实体识别与关系抽取研究提供了标准化的黄金标准基准,显著推动了生物医学文本挖掘领域的发展,并为构建miRNA调控网络奠定了坚实的数据基础。
当前挑战
该数据集面临的核心挑战在于解决生物医学文本中miRNA实体识别与关系抽取的领域难题。miRNA命名体系复杂多样,存在多种异构体、物种特异性前缀及不规则缩写,加之文献中常以模糊的上下文提及,使得基于规则的识别方法召回率有限。同时,miRNA与疾病、基因之间的关联关系高度动态且相互交织,现有方法在区分直接调控与间接关联时面临精度瓶颈,导致关系抽取性能受限。在语料库构建过程中,挑战同样严峻:需要制定详尽的标注指南以应对实体边界模糊和语义歧义问题,确保跨标注者的一致性;此外,从海量文献中筛选出包含miRNA相关信息的摘要并进行人工标注,工作量大且耗费时间,对资源有限的学术团队而言是一项艰巨任务。
常用场景
经典使用场景
miRNA数据集在生物医学自然语言处理领域中,最经典的使用场景是用于命名实体识别(NER)和命名实体消歧(NED)任务的模型训练与评估。该数据集精选了301篇Medline文献摘要,由领域专家手工标注了基因、疾病以及miRNA实体,为从非结构化文本中自动提取miRNA相关实体提供了高质量的金标准语料。研究者可借此构建和优化深度学习或基于规则的实体识别系统,从而精准定位文献中miRNA的提及,为后续关系抽取奠定坚实基础。
解决学术问题
该数据集解决了生物医学文本挖掘中缺乏标准化miRNA实体识别与关系抽取评价基准的学术难题。在miRNA研究快速增长的背景下,传统方法多依赖关键词匹配,灵敏度低且难以系统评估。miRNA语料库的构建,使得研究者能够客观比较不同算法在miRNA、基因及疾病实体识别上的性能,并推动了miRNA-疾病、miRNA-基因关系抽取方法的发展,最终助力于miRNA调控网络的系统构建与疾病机制解析。
衍生相关工作
该数据集衍生了一系列经典的生物医学文本挖掘工作,其中最具有代表性的是Bagewadi等人(2014)提出的综合方法,该方法融合了词典、正则表达式与机器学习技术,在miRNA实体识别上达到了0.94的召回率和0.93的精确率,并在miRNA-疾病和miRNA-基因关系抽取上取得了0.76的F1分数。后续研究在此基础上进一步探索了基于深度神经网络的关系抽取模型,以及跨物种miRNA提及的标准化方法,推动了miRNA文献挖掘领域的持续发展。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务