Code and Experimental Results for Hybrid CNN-Based Peer-to-Peer Credit Default Prediction
收藏资源简介:
This repository contains the source code, analysis scripts, and aggregated experimental results supporting the study entitled “Transforming Borrower Profiles into Risk Signals: A Statistically Validated Hybrid Convolutional Framework for Peer-to-Peer Credit Default Prediction.” The study investigates peer-to-peer credit default prediction using a structure-aware feature-to-image transformation combined with convolutional feature extraction and ensemble learning. The complete computational workflow includes leakage-audited preprocessing, feature selection, class-imbalance handling, feature-to-image encoding, tabular machine-learning models, standalone convolutional neural networks, CNN-based hybrid models, repeated stratified cross-validation, statistical significance testing, ablation analysis, hyperparameter sensitivity analysis, explainability analysis, and censoring-corrected temporal out-of-time validation. The repository includes Python scripts implementing the major stages of the analysis and a consolidated Excel workbook containing the experimental results of all evaluated models across the different dataset conditions and validation settings. The workbook brings together model-level performance metrics, aggregated confusion matrices, statistical comparisons, ablation results, feature-ordering analyses, hyperparameter sensitivity results, explainability analyses, and temporal out-of-time validation results, thereby facilitating transparent comparison, verification, and reproducibility of the reported findings. The reported experiments were conducted using the LendingClub accepted-loans dataset. The repository does not redistribute the original LendingClub accepted-loans dataset or any restricted third-party data. Users must obtain the original dataset from its legitimate source and comply with the applicable data-use terms and licensing conditions. The main computational components include:(1) temporal out-of-time validation;(2) preprocessing and leakage auditing;(3) feature selection;(4) data balancing using SMOTENC and downsampling;(5) feature-to-image encoding;(6) tabular classification;(7) CNN and CNN-hybrid classification;(8) repeated 10×5 stratified cross-validation;(9) Friedman and corrected resampled paired t-tests;(10) platform-risk-grade ablation analysis;(11) feature-ordering ablation using linear-rank, hierarchical-clustering, and IGTD strategies;(12) nested hyperparameter sensitivity analysis;(13) SHAP-based feature importance;(14) feature-row-aggregated Grad-CAM analysis; and(15) censoring-corrected chronological out-of-time validation. The experimental results provided in the repository cover all evaluated models and experimental conditions and are consolidated into a single Excel workbook for convenient reference and independent verification. The repository is intended to support computational reproducibility and methodological transparency of the associated study.



