BioRED
收藏资源简介:
BioRED是一个专为生物医学领域设计的丰富关系抽取数据集,由国家生物技术信息中心创建。该数据集包含600篇PubMed摘要,涉及基因、疾病、化学物质等多种实体类型及其相互关系。数据集特别之处在于对每种关系进行了新发现与背景知识的标注,以帮助算法区分这两类信息。创建过程中,研究团队采用了随机抽样和人工标注相结合的方法,确保数据的质量和代表性。BioRED数据集的应用领域广泛,旨在解决生物医学文本中信息抽取的挑战,特别是在识别新发现和避免重复信息方面。
BioRED is a rich relation extraction dataset specifically designed for the biomedical domain, created by the National Center for Biotechnology Information. This dataset comprises 600 PubMed abstracts, covering multiple entity types including genes, diseases, chemicals and their interrelations. A distinctive characteristic of this dataset is that it annotates every relationship with two categories: newly discovered findings and background knowledge, enabling algorithms to differentiate between these two types of information. During its curation, the research team employed a hybrid approach of random sampling and manual annotation to guarantee the dataset's quality and representativeness. The BioRED dataset has broad application prospects, aiming to address the challenges of information extraction from biomedical texts, especially in identifying newly discovered findings and avoiding duplicate information.




