MedSETNER: A benchmark corpus for extracting dataset names from medical scientific literature
收藏官方服务:
资源简介:
MedSETNER is the first annotated benchmark for extracting dataset names from medical and biomedical research articles. Dataset references in medical papers tend to be embedded in free-form prose, written with inconsistent nomenclature, and rarely tied to a persistent identifier. Existing dataset-mention corpora cover computer science or general science publications and do not cover medical text; that is the gap this corpus fills.
创建时间:
2026-08-05



