遇见数据集

Infection Delay Project Data

收藏
Zenodo2022-02-08 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

GitHub Repository for the project code can be found here: https://github.com/shivyucel/infection-delay-project The raw data folder contains the census and commuting data, along with their relevant shapefiles and labels. It also contains the shapefiles for the 2599 hexagons used in this analysis. The prelim_data folder contains the h3_IDs table, used to numerically identify hexagons (more info included in info file in folder). It also includes the commuting data, already preprocessed using code in the Effective Distance/SIR Model folder to filter out regions which do not have suitable data. The SIR_model_inputs folder has the mobility matrix based on the commuting data, and the population data (both already filtered to include suitable regions, same as above). The infection_delay_inputs folder contains the baseline and real mobility effective distance matrices. The mobility_data folder contains the raw mobility data from March - September 2020 (raw_mobility_data.csv), the mobility reductions across all hexagons with data around the lockdown (mobility_reduction.csv), and the mobility change (filtered_mobility_reduction.csv) used in the effective_distance_real_changes.py code, filtered for the regions which are suitable for the effective distance/SIR calculations. The result_data folder contains the tables used to generate the results, including the weighted median 10-day infection delay values for every region, the outbreak divided results, merged with income and centrality data. The final data files are 'weighted_hexagon_data.csv' and 'longford_outbreak_split_delays'. NOTE: The infection delay files take up 30+ GB of storage, and are not included here. In both mobility scenario, 2599 tables are generated, each with 2599 columns, representing the time series of every hexagon for every outbreak scenario. These are then duplicated in reorient_infection_delay_tables.py, to make them amenable to analysis. Full commuting and census data, beyond that included in the raw data folder, can be found here (https://www.metro.sp.gov.br/pesquisa-od/) and here (https://censo2010.ibge.gov.br/resultados/resumo.html), respectively.

本项目代码的GitHub仓库地址为:https://github.com/shivyucel/infection-delay-project。 原始数据文件夹(raw data folder)收录人口普查数据与通勤数据,及其配套的矢量形状文件(shapefile)与标签;同时包含本分析所用的2599个六边形网格的矢量形状文件。 预处理数据文件夹(prelim_data folder)包含用于数值标识六边形网格的`h3_IDs`表(详细说明见文件夹内的info文件),同时收录已完成预处理的通勤数据:该数据已通过「有效距离/SIR模型(Effective Distance/SIR Model)」文件夹中的代码完成过滤,仅保留具备有效数据的研究区域。 SIR模型输入文件夹(SIR_model_inputs folder)包含基于通勤数据生成的流动矩阵与人口数据,二者均已完成过滤,仅保留符合要求的研究区域,与前述预处理数据一致。 感染延迟输入文件夹(infection_delay_inputs folder)包含基准流动与真实流动对应的有效距离矩阵。 流动数据文件夹(mobility_data folder)包含2020年3月至9月的原始流动数据(raw_mobility_data.csv)、所有具备数据的六边形网格在封控期间的流动降幅数据(mobility_reduction.csv),以及用于`effective_distance_real_changes.py`脚本的流动变化数据(filtered_mobility_reduction.csv)——该数据已针对有效距离/SIR模型计算所需的研究区域完成过滤。 结果数据文件夹(result_data folder)包含用于生成研究结果的各类表格,涵盖所有区域的加权中位数10天感染延迟值、按疫情暴发时段拆分的结果,以及与收入和中心性数据(centrality data)合并后的数据集。最终数据文件为`weighted_hexagon_data.csv`与`longford_outbreak_split_delays`。 请注意:感染延迟相关文件占用存储空间超过30GB,未包含在本仓库中。在两种流动情景下,共生成2599张表格,每张表格包含2599列,对应每个暴发情景下所有六边形网格的时间序列数据;这些表格将通过`reorient_infection_delay_tables.py`脚本完成格式重定向,以适配后续分析工作。 原始数据文件夹未涵盖的完整通勤与人口普查数据,可分别从以下地址获取:https://www.metro.sp.gov.br/pesquisa-od/ 与 https://censo2010.ibge.gov.br/resultados/resumo.html。

提供机构:
Zenodo
创建时间:
2022-02-02
二维码
社区交流群
二维码
科研交流群
商业服务