semiparametric doubly-robust quantile treatment effect estimator and empirical data on environmental regulation's impact on corporate digital transformation.
收藏资源简介:
The research data encompass simulated data, simulation code, and an empirical dataset. The simulated data generate potential outcomes, observed outcomes, propensity scores, and treatment assignments under three distinct designs. The first two designs differ exclusively in their propensity score specification—linear versus nonlinear—while the third features unbalanced distributions. The simulation code implements inverse probability weighting (IPW) and doubly robust (DR) estimators for quantile treatment effects (QTE), with IPW relying on a linear parametric model and DR employing multiple machine learning algorithms for propensity score estimation, thereby validating the advantages of the semiparametric quantile DR estimator. The empirical analysis examines the causal effect of environmental regulation on corporate digital transformation, exploiting China's new energy city pilot policy as a quasi-natural experiment. Digital transformation levels are constructed through textual analysis of listed firms' annual reports. Treatment assignment is defined by region-year indicators. Confounding variables—including firm age, CEO duality, return on equity, standard audit opinion, operating revenue, total assets, and turnover ratio—serve as explanatory variables for propensity score estimation and the outcome distribution. The empirical specification employs a QTE doubly robust estimator with propensity scores estimated via random forest.
本研究数据集涵盖模拟数据、模拟代码与实证数据集。模拟数据基于三种不同的研究设计,生成潜在结果、观测结果、倾向得分(propensity score)与处理分配结果。前两种设计仅在倾向得分的设定方式上存在差异——分别为线性设定与非线性设定,而第三种设计则呈现非平衡分布特征。模拟代码实现了分位数处理效应(quantile treatment effects,QTE)的逆概率加权(inverse probability weighting,IPW)估计器与双稳健(doubly robust,DR)估计器:其中逆概率加权依托线性参数模型开展估计,双稳健则采用多种机器学习算法进行倾向得分估计,以此验证半参数分位数双稳健估计器的应用优势。实证分析聚焦于环境规制对企业数字化转型的因果效应,以中国新能源城市试点政策作为准自然实验(quasi-natural experiment)。数字化转型水平通过对上市公司年度报告的文本分析构建得到。处理分配以地区-年份标识进行定义。混淆变量涵盖企业年龄、两职合一(CEO duality)、净资产收益率(return on equity)、标准审计意见、营业收入、总资产与资产周转率,这些变量被用作倾向得分估计与结果分布建模的解释变量。实证设定采用分位数处理效应双稳健估计器,其倾向得分通过随机森林(random forest)算法完成估计。



