遇见数据集

Corpus Index for Lexical vs Semantic Retrieval Experiment in Biomedical Literature

收藏
Zenodo2026-03-20 更新2026-05-29 收录
官方服务:

资源简介:

This dataset provides the corpus index used in the controlled biomedical literature retrieval experiment reported in the associated article on lexical and semantic retrieval. The file contains article-level identifiers and metadata to support corpus reconstruction and independent verification of the study sample. Specifically, it includes: PMCID PMID Article title The indexed corpus was used to compare Bag-of-Words (BoW), TF-IDF, and embedding-based semantic retrieval methods within a fixed experimental setting. The purpose of releasing this index is to improve transparency and reproducibility while respecting copyright and licensing restrictions on the underlying full-text content. Important: The repository does not include article full text or abstracts. Researchers seeking to reproduce the analysis must obtain access to the corresponding PubMed-indexed records and, where required, full-text articles in accordance with publisher and database licensing requirements. This dataset is intended as a reconstruction aid for the controlled corpus rather than as a redistribution of the source literature.

提供机构:
Zenodo
创建时间:
2026-03-20
二维码
社区交流群
二维码
科研交流群
商业服务