Feature enhanced ensemble modeling with voting optimization for credit risk assessment
收藏资源简介:
The ChinaZJB dataset consists of 1,329 valid samples of SMEs after merging the non-financial behavioral information and soft information on credit rating with the financial information, loan information, and non-financial basic information found in the annual loan ledger data. Among them, 108 SMEs have default records, while 1,221 SMEs have no default records, resulting in an imbalanced ratio of approximately 1:11.Five datasets from the UC Irvine (UCI) machine-learning repository, that is, the Polish 1, Polish 2, Polish 3 , Australian, and Taiwan credit datasets, were used for robustness checks.
ChinaZJB数据集整合了年度贷款台账数据中的金融信息、贷款信息、非金融基础信息,以及信用评级相关的非金融行为信息与软信息,最终得到1329条有效中小企业(Small and Medium Enterprises,SMEs)样本。其中,108家中小企业存在违约记录,1221家中小企业无违约记录,样本类别不平衡比例约为1:11。本研究同时采用加州大学欧文分校(University of California, Irvine,UCI)机器学习库中的5个公开数据集,即波兰1、波兰2、波兰3、澳大利亚信用数据集以及台湾信用数据集,用于模型的稳健性检验。




