遇见数据集

Monte Carlo Simulation Data and Code for Threshold Policies in Educational Early Warning Systems

收藏
Zenodo2025-12-05 更新2026-05-26 收录
官方服务:

资源简介:

This record contains the complete data, code, and figures for a Monte Carlo simulation study on threshold policies in educational early warning systems (EWS). The simulations examine how different ways of converting risk scores into alerts—such as workload caps, recall-oriented rules, group-specific thresholds, cost-sensitive policies, and precision-first rules—shape the joint behaviour of workload, predictive performance, and group-based fairness when flagging students at risk of dropout or academic failure. Rather than optimising prediction models in isolation, the study focuses on the decision layer where continuous risk scores are turned into discrete alerts under resource constraints. The “true world” is a stylised but realistic logistic data-generating process with two time points (T1, T2), two continuous indicator families (reverse-coded readiness and emotional distress), and two student groups with different base rates and risk profiles. Four design factors are fully crossed: the overall base rate of at-risk students (0.20, 0.30, 0.40), group separation on the indicators (d = 0.00, 0.30, 0.60), total sample size (N = 600, 1200), and whether the prediction models include the group indicator G or omit it. Crossing these factors yields 36 design conditions. For each condition, two logistic prediction models (T1-only vs. T1+T2) are fitted on a 70% training split and evaluated on a 30% test split, and five threshold-policy families are derived on the training data and then applied to test-set probabilities. Each cell is replicated 10,000 times with independent seeds. The deposit is organised into a fig folder with all figures (workload–performance, fairness–workload, cost–fairness, and sensitivity plots), an R code folder containing the main simulation script (threshold_sim.R), a Result/Full Data folder with per-condition CSV files in both replication-level (“long”) and Monte Carlo summary (“summ”) formats plus two combined “36-condition” tables, and a threshold folder with an RStudio project (threshold.Rproj) and renv files to recreate the original package environment. Together, these materials allow users to reproduce all results, inspect the full distribution of metrics across replications, and reuse or extend the simulation framework for other early-warning or algorithmic-fairness studies.

提供机构:
Zenodo
创建时间:
2025-12-05
二维码
社区交流群
二维码
科研交流群
商业服务