Rural E-commerce, Ecological Value Conversion, and Common Prosperity: Cleaned County-Year Panel (2014–2022)
收藏资源简介:
The panel covers 2{,}725 Chinese counties across 30 provincial units from 2014 to 2022 and supports three nested analysis samples used in the paper. All variables have been harmonised to a single county identifier scheme, Winsorised at the 1% level where appropriate, and documented for reproducibility. When citing the underlying study, please also cite the manuscript that uses this panel. ## File Inventory The archive `cleaned_data.zip` contains 23 files organised into five functional groups. ### Master Panel | File | Rows | Unit | Years | Purpose | |---|---:|---|---|---| | `master_panel_county_2014_2022.csv` | 24,525 | county-year | 2014–2022 | Primary panel with 44 curated analytic variables | | `master_panel_full.csv` | 24,525 | county-year | 2014–2022 | Wide version with all merged fields (87 columns) | ### Analytic Samples | File | Rows | Counties | Notes | |---|---:|---:|---| | `baseline_sample.csv` / `baseline_sample.dta` | 9,247 | 1,600 | Listwise deletion on Y/X/Q/controls, 1% Winsor, ready for fixed-effect regressions | | `did_sample.csv` / `did_sample.dta` | 12,498 | 1,929 | Loose sample preserving post-treatment years for staggered DID | ### Component Sub-Tables | File | Rows | Unit | Years | |---|---:|---|---| | `taobao_prov_panel.csv` | 279 | province | 2014–2022 | | `taobao_city_panel.csv` | 1,701 | prefecture | 2014–2022 | | `taobao_county_panel.csv` | 5,994 | county | 2014–2022 | | `taobao_unmatched.csv` | 189 | county name | — | | `yearbook_county_panel_2014_2022.csv` | 24,525 | county-year | 2014–2022 | | `county131_panel.csv` | 51,742 | county-year | 2000–2019 | | `m_county_panel.csv` | 24,525 | county-year | 2014–2022 | | `gini_county_panel.csv` | 69,048 | county-year | 2000–2023 | | `dfiic_county_panel.csv` | 26,091 | county-year | 2014–2023 | | `did_digital_village_panel.csv` | 64,345 | county-year | 2000–2023 | | `gtfp_city_panel.csv` | 5,076 | city-year | 2006–2023 | | `rural_revit_city_panel.csv` | 10,104 | city-year | 2000–2023 | | `internet_prov_panel.csv` | 403 | province-year | 2011–2023 | | `internet_city_panel.csv` | 6,196 | city-year | 2003–2022 | | `common_prosperity_prov_panel.csv` | 682 | province-year | 2000–2021 | ### Classification Labels | File | Rows | Purpose | |---|---:|---| | `county_groupings.csv` | 2,725 | Time-invariant county labels for heterogeneity analysis: region (east/middle/west), staple-grain vs cash-crop, eco-sensitive indicators | ### Quality Report | File | Purpose | |---|---| | `QC_report.md` | Missing-rate diagnostic, per-year and per-province coverage for the master panel | ## Variable Naming Conventions Variables in the analytic samples follow a consistent prefix system. Variables starting with `y_` are outcomes. Variables starting with `x_` are treatment measures of Taobao village density. Variables starting with `m_` are mediators related to agricultural land-based value conversion efficiency. Variables starting with `q_` are moderators derived from the Peking University Digital Financial Inclusion Index. Variables starting with `ctrl_` are control variables. The fields `treat`, `post`, and `did` encode the staggered treatment indicators of the digital-village pilot launched in 2020. The primary outcome variables are `y_sen_welfare` (Sen social welfare function on rural income and Gini), `y_rural_inc_log` (natural logarithm of rural per-capita disposable income), and `y_gini` (nightlight-based county Gini). The primary treatment is `x_taobao_density` (Taobao villages per ten thousand residents) with its quadratic `x_taobao_density_sq`. The primary mediator is `m_primary`, defined as the natural logarithm of agricultural output value divided by administrative land area. ## Raw Data Provenance The panel integrates administrative and platform data from the following sources. County-level statistics come from the *China County Statistical Yearbook* compiled through the 2022 edition. Taobao village counts come from the annual rosters published by the AliResearch Institute between 2014 and 2022, with the 2022 roster providing a twelve-digit village code that anchors cross-year county identification. Digital financial inclusion indicators are from the Peking University Digital Financial Inclusion Index (2014–2023 version). Nightlight-based Gini coefficients follow the county-level construction by Chen et al. (2015). The digital-village pilot roster was obtained from the Ministry of Agriculture and Rural Affairs and the Cyberspace Administration of China announcements. Prefecture-level green total factor productivity is from published panel estimates covering 2006–2023. ## Reproducibility The cleaning pipeline is available as a set of Python scripts in the accompanying repository. The pipeline runs end to end with pandas, numpy, and openpyxl. Statistical analyses use pyfixest (primary), linearmodels, statsmodels, and the differences package for staggered DID. The ordered build is documented in the main paper's online appendix. Stata dta exports are provided for users who prefer to re-estimate specifications in Stata. The character encoding of all CSV files is UTF-8 with BOM, which renders correctly in Microsoft Excel. ## Known Limitations Roughly 58 percent of county-year observations in `m_primary` carry missing values, a result of coverage gaps in the underlying agricultural output series. Multiple imputation is used in the paper as a robustness check. The matching between Taobao village names and national county codes succeeds for 96.9 percent of village-year rows, with 189 unmatched rows covering 21 distinct county names preserved in `taobao_unmatched.csv` for transparency. The baseline sample loses 2022 observations because a small number of control variables are not yet released for that year in the current yearbook edition, and users who require a strict 2022 balanced panel should consult the DID sample instead. County-level gross ecosystem product measures are not yet publicly available for a sufficiently broad set of counties, and the mediator used here relies on administrative agricultural output as a proxy. ## License The data are released under the Creative Commons Attribution 4.0 International license (CC-BY 4.0), which permits reuse and redistribution provided attribution is given to the authors. Downstream users should comply with the license terms of the original raw sources where onward redistribution is concerned.
本面板数据集涵盖2014至2022年中国30个省级行政区的2725个县域,支持论文中使用的三类嵌套分析样本。所有变量均已统一为单一县域识别体系,必要时已进行1%水平的温塞标准化(Winsorization)处理,并附带可复现性文档。 当引用本基础研究时,请同时引用使用该面板的相关论文。 ## 文件清单 归档文件`cleaned_data.zip`包含23个文件,分为5个功能组。 ### 核心面板(Master Panel) | 文件 | 行数 | 分析单元 | 年份范围 | 用途 | |---|---:|---|---|---| | `master_panel_county_2014_2022.csv` | 24,525 | 县域-年份 | 2014–2022 | 包含44个经过筛选的分析变量的核心面板 | | `master_panel_full.csv` | 24,525 | 县域-年份 | 2014–2022 | 包含所有合并字段的宽格式面板(共87列) | ### 分析样本(Analytic Samples) | 文件 | 行数 | 县域数量 | 说明 | |---|---:|---:|---| | `baseline_sample.csv` / `baseline_sample.dta` | 9,247 | 1600 | 对结果变量、处理变量、调节变量及控制变量进行列表式删除,已完成1%温塞标准化,可直接用于固定效应回归 | | `did_sample.csv` / `did_sample.dta` | 12,498 | 1929 | 宽松样本,保留渐进式双重差分(staggered Difference-in-Differences, DID)所需的后处理年份数据 | ### 组件子表(Component Sub-Tables) | 文件 | 行数 | 分析单元 | 年份范围 | |---|---:|---|---| | `taobao_prov_panel.csv` | 279 | 省级 | 2014–2022 | | `taobao_city_panel.csv` | 1,701 | 地级市 | 2014–2022 | | `taobao_county_panel.csv` | 5,994 | 县域 | 2014–2022 | | `taobao_unmatched.csv` | 189 | 县域名称 | — | | `yearbook_county_panel_2014_2022.csv` | 24,525 | 县域-年份 | 2014–2022 | | `county131_panel.csv` | 51,742 | 县域-年份 | 2000–2019 | | `m_county_panel.csv` | 24,525 | 县域-年份 | 2014–2022 | | `gini_county_panel.csv` | 69,048 | 县域-年份 | 2000–2023 | | `dfiic_county_panel.csv` | 26,091 | 县域-年份 | 2014–2023 | | `did_digital_village_panel.csv` | 64,345 | 县域-年份 | 2000–2023 | | `gtfp_city_panel.csv` | 5,076 | 城市-年份 | 2006–2023 | | `rural_revit_city_panel.csv` | 10,104 | 城市-年份 | 2000–2023 | | `internet_prov_panel.csv` | 403 | 省级-年份 | 2011–2023 | | `internet_city_panel.csv` | 6,196 | 城市-年份 | 2003–2022 | | `common_prosperity_prov_panel.csv` | 682 | 省级-年份 | 2000–2021 | ### 分类标签(Classification Labels) | 文件 | 行数 | 用途 | |---|---:|---| | `county_groupings.csv` | 2,725 | 用于异质性分析的时不变县域标签:包括区域划分(东/中/西部)、粮食主产区vs经济作物区、生态敏感指标 | ### 质量报告(Quality Report) | 文件 | 用途 | |---|---| | `QC_report.md` | 核心面板的缺失率诊断报告,包含分年度、分省级行政区的覆盖情况 | ## 变量命名规范 分析样本中的变量采用统一的前缀体系:以`y_`开头的变量为结果变量;以`x_`开头的变量为淘宝村密度的处理变量;以`m_`开头的变量为与农用地价值转化效率相关的中介变量;以`q_`开头的变量为来自北京大学数字普惠金融指数(Peking University Digital Financial Inclusion Index)的调节变量;以`ctrl_`开头的变量为控制变量。字段`treat`、`post`和`did`用于编码2020年启动的数字乡村试点的渐进式处理指标。 核心结果变量包括`y_sen_welfare`(基于农村收入与基尼系数的森社会福利函数)、`y_rural_inc_log`(农村人均可支配收入对数)和`y_gini`(基于夜光数据的县域基尼系数)。核心处理变量为`x_taobao_density`(每万居民对应的淘宝村数量)及其二次项`x_taobao_density_sq`。核心中介变量为`m_primary`,定义为农业总产值除以行政区域面积的自然对数。 ## 原始数据来源 本面板整合了来自以下渠道的行政与平台数据:县域层面统计数据来自截至2022版的《中国县域统计年鉴》;淘宝村数量来自阿里研究院(AliResearch Institute)2014至2022年发布的年度名录,其中2022年名录提供了12位村级代码,用于锚定跨年度县域识别;数字普惠金融指数来自北京大学数字普惠金融指数(2014–2023版);基于夜光数据的基尼系数参考了Chen等人(2015)的县域构建方法;数字乡村试点名录来自农业农村部(Ministry of Agriculture and Rural Affairs)与国家互联网信息办公室(Cyberspace Administration of China)的公告;地级市层面绿色全要素生产率(green total factor productivity)数据来自2006–2023年已发布的面板估计结果。 ## 可复现性说明 数据清洗流程以Python脚本形式集成于配套开源仓库中,流程全程依赖pandas、numpy与openpyxl库实现。统计分析主要使用pyfixest,同时支持linearmodels、statsmodels及用于渐进式DID的differences包。有序构建流程已在主论文的在线附录中说明。同时提供Stata数据文件(Stata dta)导出格式,供偏好使用Stata进行模型重估计的用户使用。所有CSV文件均采用带BOM的UTF-8编码,可在Microsoft Excel中正常读取。 ## 已知局限性 `m_primary`变量中约58%的县域-年份观测值存在缺失,这源于原始农业产出序列的覆盖缺口,论文中采用多重插补(multiple imputation)作为稳健性检验。淘宝村名称与全国县域代码的匹配成功率为96.9%,共189条未匹配行对应21个不同县域名称,已保存至`taobao_unmatched.csv`以保证透明度。基线样本剔除了2022年的观测值,原因是当年部分控制变量尚未在最新版年鉴中发布,若需严格的2022年平衡面板(balanced panel),请参考DID样本。县域层面生态系统生产总值(gross ecosystem product)尚未面向足够广泛的县域公开,本研究使用的中介变量以行政农业产出作为代理变量。 ## 许可协议 本数据集采用知识共享署名4.0国际许可协议(Creative Commons Attribution 4.0 International license,CC-BY 4.0)发布,允许在注明原作者的前提下进行重用与再分发。下游用户在进行二次再分发时,需遵守原始数据源的许可条款。



