遇见数据集

Scientific literature on the nexus of climate change and health

收藏
Zenodo2026-04-08 更新2026-05-26 收录
官方服务:

资源简介:

Systematic map of scientific literature on climate change and health and attributable impacts This repository contains the raw data used in the Lancet Countdown 2026 (5.3.1/5.3.2). This dataset is also underlying the climateliterature.org website (https://lithub.pik-potsdam.de/#/project/healthmap_2026). As part of the Pathfinder and DESTinY projects funded by the Wellcome Trust (UK) and in the context of the Lancet Countdown, we plan to regularly update this dataset as new data becomes available.Note, that these files need to be joined and filtered in the correct way. It is highly recommended to use the accompanying Python package to read the data and take the analysis notebooks as a reference of how to correctly derive numbers:https://gitlab.pik-potsdam.de/mcc-apsis/living-evidence-maps/climate-health-map/-/blob/v1.0/Notebooks/Global_git.ipynb In general, it is always important to not simply count the number of rows, but the number of unique item_ids in the relevant set. Files items.csv: This is the main file listing all records included in this export. In this particular case, this does not contain all records matching the search query, but only those that are actually classified as relevant. Because of potential copyright issues we removed the abstracts. affiliations.csv: This table can be joined on the item_id columns to the items file. It contains all the authorship affiliation information provided by OpenAlex. We don't have authorship information for all records or maybe the affiliation could not be explicitly matched to a specific country. places.csv: This table can be joined on the item_id columns to the items file. It contains all locations identified by mordecai3 in the text. This is not cleaned or filtered, but the package mentioned above has the correct data loading and filtering methods. We share the full raw data here in case classifications.csv: This table can be joined on the item_id columns to the items file. Each column corresponds to a label classified using classifier and the cell contains the respective confidence score. Some columns contain topic model scores. Not all cells are filled, for example because the topic score is small or an item was not classified further because it was marked as irrelevant before. Typically, one can assume a score >= 0.5 to mean that the label is present. Labels are grouped, for example climate category (cat) has three columns: mitigation, adaptation, and impacts. We use the "multi-label" logic in most cases, that means from the same group, multiple labels a might apply. However, in some cases, it might be preferred to choose the label with the highest score instead. In that case, one should check that at least one of the scores is >= 0.5. In either case, it is important to check that the correct reference total is chosen when normalising (for example calculating the proportion of records per year on a certain climate category). For specifics, it is hightly recommended to review how filters are applied and counts produced in the Notebooks in the accompanying python package. The Excel sheets contain aggregated counts for both indicators and ways of counting (primary class or multi-class) References When using this data, please make sure to cite the dataset. You may also cite the Lancet Countdown (Global or regional if using a subset). The methodology and most human annotations are derived from Berrang-Ford L, Sietsma AJ, Callaghan M, et al. Systematic mapping of global research on climate and health: a machine learning review. Lancet Planet Health 2021; 5: e514–25. Callaghan M, Schleussner CF, Nath S, et al. Machine-learning-based evidence and attribution mapping of 100,000 climate impact studies. Nat Clim Chang 2021; 11: 966–72.

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务