Mining Local Government Behavior: A News-Based Panel Dataset of Credit Governance in China (2000–2025) Generated via Large Language Models
收藏资源简介:
China’s Social Credit System represents one of the most significant experiments in digital governance globally, yet empirical research has long been constrained by the lack of micro-level, high-frequency execution data. Existing datasets typically rely on macroscopic statistics or static policy documents, failing to capture the dynamic attention allocation and diverse policy instruments employed by local governments. To address this gap, we present a large-scale spatiotemporal panel dataset of local government credit governance in China, spanning from 2000 to 2025. The dataset is derived from a corpus of over 5 million news articles collected from official government portals and state media across 31 provinces, 306 prefectures, and 469 counties. We employed a novel Dual-Track Information Extraction Strategy using Large Language Models (LLMs): a Silver Track utilizing the Qwen-7B model for macro-classification of governance domains and sentiments, and a Gold Track utilizing the Qwen-72B model for the micro-extraction of specific regulatory mechanisms and lifecycle phases. Furthermore, we bridged the gap between unstructured text and quantitative social science by aligning the news-derived indicators with long-term socio-economic statistics (2000–2024) via standardized administrative division codes. This dataset provides a comprehensive, granular, and analysis-ready resource for researchers to evaluate policy effectiveness, analyze government behavior patterns, and explore regional heterogeneity in China’s modernization of governance.



