A Longitudinal Dataset for Machine Learning-Based Student Retention Analysis in Bangladeshi STEM Higher Education
收藏资源简介:
Educational data mining in Bangladesh has been limited by the absence of localized, longitudinal datasets suitable for machine learning-based retention analysis. This dataset provides anonymized academic records of 3,762 Computer Science and Engineering (CSE) students from Leading University, Sylhet, Bangladesh, spanning 16 years from Spring 2010 to Spring 2025. The dataset includes 11 attributes covering student demographics, semester-wise academic performance (CGPA, credits completed), admission cohort, and an engineered retention status label (Graduate/Eligible, In Progress, Dropout) derived through cohort-based classification rules. Baseline machine learning benchmarks are provided using six standard classifiers with 5-fold cross-validation, achieving a predictive accuracy of 77.7% and an AUC of 0.82 for dropout prediction. Statistical analysis also confirms a significant rise in female student enrollment and academic performance over the 16-year period (p < 0.001). This resource addresses a critical gap in localized longitudinal datasets for educational data mining in developing-country STEM contexts. It is intended to support researchers and policymakers in developing Early Warning Systems (EWS), predictive retention models, and data-driven strategies to strengthen higher education in Bangladesh. This dataset accompanies the paper "A Longitudinal Dataset for Machine Learning-Based Student Retention Analysis in Bangladeshi STEM Higher Education," accepted at the 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN), 16–18 April 2026, Chittagong, Bangladesh.



