遇见数据集

BIOMAT-CellNER: A Biomaterials Domain-Specific Corpus for Named Entity Recognition of Cell Mentions

收藏
Zenodo2025-04-28 更新2026-05-26 收录
官方服务:

资源简介:

BIOMAT-CellNER Corpus BIOMAT-CellNER stands for BIOMATerials Cell Named Entity Recognition. It is a corpus developed within the scope of the Horizon Europe BIOMATDB project to support the extraction and classification of cell entity mentions from the scientific literature in the biomaterials domain. It focuses on the annotation of cell mentions, including both cell types (e.g., fibroblasts, stem cells) and cell lines (e.g., HeLa, MC3T3), particularly those that are studied in interaction with biomaterials in experimental settings. The corpus was created through a collaborative effort involving domain experts, who were tasked with the establishment of comprehensive and accurate annotation guidelines for the manual annotation of the final gold standard corpus. To ensure domain relevance and terminological coverage, PubMed abstracts were carefully selected based on relevant MeSH (Medical Subject Headings) categories associated with biomaterials, tissue engineering, and related fields. The abstracts were then manually annotated according to the predefined rules in the annotation guidelines. The BIOMAT-CellNER corpus is one of four developed within the project and is divided into three subsets: a training set (750 documents), a test set (150 documents), and a validation set (100 documents), available in multiple formats, including brat, CSV and CoNLL. This corpus is part of a broader initiative to support the development of an advanced, searchable biomaterials database with integrated analytical tools and digital advisors. It is also intended for use in training Named Entity Recognition (NER) models, enabling the automatic identification and extraction of cell mentions relevant to biomaterials research. Resources Project Website Biomaterials Marketplace Biomaterials Database

BIOMAT-CellNER语料库 BIOMAT-CellNER即BIOMATerials Cell Named Entity Recognition(生物材料细胞命名实体识别)语料库。它是在欧盟地平线欧洲(Horizon Europe)BIOMATDB项目框架下开发的语料库,旨在支持从生物材料领域的科学文献中提取并分类细胞实体提及内容。 该语料库聚焦细胞提及的标注工作,涵盖细胞类型(如成纤维细胞、干细胞)与细胞系(如HeLa、MC3T3),尤其针对实验环境中与生物材料相互作用的研究细胞。 本语料库由领域专家协作完成,专家们制定了全面且精准的标注规范,用于手动标注最终的金标准语料库。为确保领域相关性与术语覆盖度,研究人员基于与生物材料、组织工程及相关领域关联的医学主题词表(MeSH,Medical Subject Headings)类别,精心筛选了PubMed摘要。随后依据标注规范中的预定义规则,对这些摘要进行手动标注。 BIOMAT-CellNER语料库是该项目开发的四个语料库之一,被划分为三个子集:训练集(750份文档)、测试集(150份文档)与验证集(100份文档),并提供brat、CSV及CoNLL等多种格式。 该语料库是一项更广泛倡议的组成部分,该倡议旨在开发一款集成分析工具与数字顾问的高级可搜索生物材料数据库。同时,它也可用于训练命名实体识别(Named Entity Recognition,NER)模型,实现对生物材料研究相关的细胞提及内容的自动识别与提取。 资源 项目官网 生物材料交易市场 生物材料数据库

提供机构:
Zenodo
创建时间:
2025-04-25
二维码
社区交流群
二维码
科研交流群
商业服务