Gold-standard annotation sets and reproducible NLP pipeline for recovering PRO evidence from cataract-surgery literature (2000-2025)
收藏资源简介:
This repository contains the gold-standard annotation sets and analytical scripts used in the study "Closing the Patient-Reported Outcome Measurement Gap in Cataract Care: A Natural Language Processing Evidence Synthesis, 2000–2025".The dataset includes: - 200 Subject-Verb-Object (SVO) expert-adjudicated extractions. - 120 manually classified photic-phenomenon mentions (glare, halo, starburst). - 268 abstracts evaluated for named PROMs instruments.Also included are the Python scripts for Latent Dirichlet Allocation (LDA) topic modeling, syntactic dependency parsing via spaCy, and clause-level context classification.Note: Raw bibliographic records from Web of Science and Scopus are not included due to database licensing restrictions. The search strategies required to reproduce the corpus are detailed in the manuscript's Methods section.



