GreenInformationFactory – WP1/D1.2 Literature Corpus: Tidy Data, Hotspot Analytics & ML-Assisted Coding Models
收藏资源简介:
Derived FAIR release generated by the GreenInformationFactory pipeline (gif.literature / gif.lit_analytics / gif.lit_ml) from the BioFairNet WP1/D1.2 literature datasets. Contains: (1) tidy corpus tables (papers.csv: 366 papers with English snake_case columns; codes_long.csv: manual codes in long format; papers_coded.csv: joined wide table), (2) descriptive hotspot analytics (code frequencies by sector, country mentions, geographic levels, publication years, barrier-driver co-occurrence; tables and figures), and (3) TF-IDF text-classification models trained on the manual codes for pre-coding future literature batches (literature_coder.pkl with cross-validated evaluation tables; screening aid, not a replacement for manual coding). Source data (CC-BY-4.0) by Guerreschi, Lomuscio & Albanese: full list 10.5281/zenodo.20743706, codebook 10.5281/zenodo.20744025. Pipeline software: 10.5281/zenodo.18428335.



