Manually Validated Thematic and Rhetorical Annotations for Health Misinformation Corpora
收藏资源简介:
This dataset contains de-identified thematic and rhetorical annotations for a stratified qualitative sample of 482 records drawn from three health-misinformation corpora: COVID-19-FNIR, Constraint, and the Monkeypox misinformation: Twitter dataset. The annotations were developed to support qualitative inspection and triangulation of quantitative linguistic, psychological, engagement, and classification findings. The sample comprises 152 COVID-19-FNIR records, 214 Constraint records, and 116 Monkeypox records. The Monkeypox subcorpus was stratified by binary class and observed engagement, with high engagement defined using the empirical 90th-percentile threshold of the composite engagement measure. Each record includes a primary and optional secondary thematic category, a primary and optional secondary rhetorical frame, a confidence rating, a borderline-case designation, and an annotation note. All 482 records were manually reviewed and validated by the creator. The release contains de-identified source-record identifiers, annotation labels, the coding framework, operational definitions, sampling documentation, a data dictionary, and source-file checksums. Original post text, usernames, profile information, and direct links are not redistributed. Researchers must obtain authorised copies of the original datasets and comply with their respective licences and platform terms. The annotations were produced in support of the doctoral thesis, Psychologically Informed Machine Learning for Health Misinformation Detection and Propagation in Online Social Networks (OSNs), undertaken in the Department of Computing and Mathematics at Manchester Metropolitan University.



