遇见数据集

Pharmacokinetic Entity Normalisation Corpora for Biomedical Literature Text and Tables

收藏
Zenodo2026-07-26 更新2026-08-02 收录
官方服务:

资源简介:

This repository contains expert-annotated datasets for pharmacokinetic (PK) entity normalisation in biomedical literature. The datasets support the development and evaluation of methods that map textual mentions of PK parameters to standardised concepts in an expanded PK parameter ontology. Two complementary corpora are provided: Sentence corpus: PK parameter mentions identified in sentences sampled from PubMed abstracts. Table corpus: PK parameter mentions identified in cells sampled from PK-relevant tables in PubMed Central Open Access articles. Each corpus is divided into predefined training, validation, and test sets. Every record represents one candidate PK mention and includes the source text, character-level entity span, tokenisation information, and a gold-standard ontology label. Sentence records include bibliographic metadata such as the PMID and PubMed URL. Table records include the complete source table in HTML format, together with its caption, footer, table identifier, and target-cell coordinates. Candidate mentions were independently annotated by two pharmacometricians using an expanded PK ontology, with disagreements resolved by a third reviewer. Mentions that did not correspond to an available ontology concept, including non-PK entities or out-of-scope parameters, were assigned the ontology’s NIL label. The annotation process and resulting corpora were developed for evaluating heuristic, representation-based, and large-language-model-assisted approaches to PK entity normalisation. The PK ontology and annotation guidelines for the sentence and table corpus are included. Please see the included README for further details.

提供机构:
Zenodo
创建时间:
2026-07-26
二维码
社区交流群
二维码
科研交流群
商业服务