遇见数据集

BIOMAT-CellNER: A Biomaterials Domain-Specific Corpus for Named Entity Recognition of Cell Mentions

收藏
Zenodo2025-04-28 更新2026-05-26 收录
官方服务:

资源简介:

BIOMAT-CellNER Corpus BIOMAT-CellNER stands for BIOMATerials Cell Named Entity Recognition. It is a corpus developed within the scope of the Horizon Europe BIOMATDB project to support the extraction and classification of cell entity mentions from the scientific literature in the biomaterials domain. It focuses on the annotation of cell mentions, including both cell types (e.g., fibroblasts, stem cells) and cell lines (e.g., HeLa, MC3T3), particularly those that are studied in interaction with biomaterials in experimental settings. The corpus was created through a collaborative effort involving domain experts, who were tasked with the establishment of comprehensive and accurate annotation guidelines for the manual annotation of the final gold standard corpus. To ensure domain relevance and terminological coverage, PubMed abstracts were carefully selected based on relevant MeSH (Medical Subject Headings) categories associated with biomaterials, tissue engineering, and related fields. The abstracts were then manually annotated according to the predefined rules in the annotation guidelines. The BIOMAT-CellNER corpus is one of four developed within the project and is divided into three subsets: a training set (750 documents), a test set (150 documents), and a validation set (100 documents), available in multiple formats, including brat, CSV and CoNLL. This corpus is part of a broader initiative to support the development of an advanced, searchable biomaterials database with integrated analytical tools and digital advisors. It is also intended for use in training Named Entity Recognition (NER) models, enabling the automatic identification and extraction of cell mentions relevant to biomaterials research. Resources Project Website Biomaterials Marketplace Biomaterials Database

BIOMAT-CellNER语料库 BIOMAT-CellNER是BIOMATerials Cell Named Entity Recognition的缩写,即生物材料细胞命名实体识别(Named Entity Recognition, NER)。该语料库由欧盟地平线欧洲(Horizon Europe)框架下的BIOMATDB项目研发,用于从生物材料领域的学术文献中提取并分类细胞实体提及内容。其核心聚焦于细胞提及的标注工作,涵盖细胞类型(如成纤维细胞、干细胞)与细胞系(如HeLa、MC3T3),特别是那些在实验场景中与生物材料开展相互作用研究的细胞。本语料库由领域专家协作构建,专家团队需制定全面且精准的标注规范,用于最终金标准语料库的人工标注工作。为保障领域相关性与术语覆盖范围,研究人员基于与生物材料、组织工程及相关领域匹配的医学主题词表(Medical Subject Headings, MeSH)类别,从PubMed数据库中精心筛选摘要文本,随后依据标注规范中的预定义规则对筛选出的摘要开展人工标注。 BIOMAT-CellNER语料库是该项目研发的四个语料库之一,被划分为三个子集:训练集(750篇文档)、测试集(150篇文档)与验证集(100篇文档),支持brat、CSV及CoNLL等多种格式。 本语料库隶属于一项更宏大的倡议,该倡议旨在打造一款集成分析工具与数字顾问的高级可搜索生物材料数据库。同时,它也可用于训练命名实体识别(Named Entity Recognition, NER)模型,实现与生物材料研究相关的细胞提及内容的自动识别与提取。 资源 项目官网 生物材料市场 生物材料数据库

提供机构:
Zenodo
创建时间:
2025-04-25
二维码
社区交流群
二维码
科研交流群
商业服务