遇见数据集

HDN-Rare

收藏
Zenodo2026-07-05 更新2026-08-01 收录
官方服务:

资源简介:

The dataset primarily consists of the underlying literature corpus, dual-format NER annotation data, an entity-linking knowledge base, and foundational ontology files. Specifically, the pubmed_rare_disease_abstracts_2025_9.json file stores 4,783 filtered PubMed abstracts in JSON list format. Each record contains complete metadata, including the PMID, title, full abstract text, publication year, and journal name. To accommodate different modeling paradigms, the annotation data are organized into two independent formats. Data intended for generative tasks are stored in the HDN-Rare_Entity_List folder, which contains train.json, dev.json, and test.json. Each sentence-level instance provides the original text (text), a source-tracking identifier (orig_id), and an entity list containing entity type information (entities). Data intended for sequence labeling tasks are stored in the HDN-Rare_Mention_Span folder, which contains files such as train.jsonlines. This format provides a tokenized token list (tokens) together with entity span information (entity_mentions), including exact start and end offsets (start, end).For auxiliary knowledge resources, the HDN-Rare_Linked folder contains seven independent JSON dictionary files corresponding to the seven core entity categories. These files record mappings between entity mentions and standardized database identifiers, such as NCBI Gene IDs. Finally, the ORDO.csv file comprehensively contains rare disease ontology information derived from Orphanet and serves as the core dictionary for defining Rare Disease entities in this dataset.

提供机构:
Zenodo
创建时间:
2026-07-05
二维码
社区交流群
二维码
科研交流群
商业服务