Corpus Index for Lexical vs Semantic Retrieval Experiment in Biomedical Literature
收藏资源简介:
This dataset provides the corpus index used in the controlled biomedical literature retrieval experiment reported in the associated article on lexical and semantic retrieval. The file contains article-level identifiers and metadata to support corpus reconstruction and independent verification of the study sample. Specifically, it includes: PMCID PMID Article title The indexed corpus was used to compare Bag-of-Words (BoW), TF-IDF, and embedding-based semantic retrieval methods within a fixed experimental setting. The purpose of releasing this index is to improve transparency and reproducibility while respecting copyright and licensing restrictions on the underlying full-text content. Important: The repository does not include article full text or abstracts. Researchers seeking to reproduce the analysis must obtain access to the corresponding PubMed-indexed records and, where required, full-text articles in accordance with publisher and database licensing requirements. This dataset is intended as a reconstruction aid for the controlled corpus rather than as a redistribution of the source literature.



