Curated Biomedical Abstracts Dataset - Provenance-Verified, Deduplicated, Refined Labels (v1.0)
收藏资源简介:
Title: Biomedical Abstracts Corpus — Provenance-Verified, Deduplicated, Refined Labels (v1.0) DOI: 10.5281/zenodo.17229456 Date created: 2025-09-29 Creators: Idris Babalola, Adewale Alex Adegoke, Peter Adebayo Odesola Contact: eidreiz01@gmail.com ====================================Dataset overviewThis deposit provides a curated biomedical abstracts corpus derived from a widely used Kaggle community dataset. The aim is a reusable, provenance-verified resource for text-classification tasks across major disease areas (e.g., cardiovascular, digestive system, neoplasms, nervous system). ==================================== Scope & counts Total records in source: 14,438 Duplicates/near-duplicates removed: 6,140 Additional filtering: removed “General Pathological Conditions” class and short abstracts (length-based threshold) Final unique records: 5,901 Final label set: Cardiovascular Diseases, Digestive System Diseases, Neoplasms, Nervous System Diseases ====================================Provenance verification (PubMed spot-check)A random sample of records was cross-referenced against PubMed. 99/100 sampled records returned PMIDs, supporting source traceability and dataset credibility. (See docs pubmed_spotcheck_results.csv included.) Note: Code is archived separately and will also be made available on GitHub soon. ==================================== Methods Deduplication: combined text-similarity and metadata to remove duplicates. Label design: normalized disease labels and retired the broad “General Pathological Conditions” class after six diagnostics checks(e.g i. lexical overlap, ii. unsupervised separability based on TF-IDF(1-2) class centroids, iii. intra-class cohesion, distinctiveness estimated via log-odds with an informative Dirichlet prior), to reduce semantic bleed-over that hindered class separability. Short-abstract filter: removed entries below a minimum length threshold to improve training quality. Provenance: PubMed ID recovery via title/metadata matching; PMIDs recorded where found. ====================================Intended use Biomedical NLP: supervised/unsupervised text classification, benchmarking, Reliability & Model uncertainty, LLM use case. Curation & reproducibility: exemplars for deduplication, provenance checks, and label-taxonomy repair. ====================================Licensing This derived dataset (metadata, labels, keys): CC BY-SA 3.0. Original dataset license: CC BY-SA 3.0 (see https://github.com/sebischair/Medical-Abstracts-TC-Corpus & https://huggingface.co/datasets/TimSchopf/medical_abstracts). Redistribution is permitted under ShareAlike terms; please attribute both the original creators and this curated release. Abstract text itself may be subject to third-party rights. Users should respect source terms when retrieving full text. ====================================Funding & competing interestsThis work received no funding. The authors declare no conflicts of interest. The curation and validation were conducted independently and are not endorsed by our current or past affiliations. ==================================== Acknowledgments & original sourceThis corpus is derived from the original community dataset created by https://www.kaggle.com/datasets/chaitanyakck/medical-text. We thank the original creators for making their resource available under CC BY-SA 3.0.



