PREDACTBENCH
收藏资源简介:
PREDACTBENCH是一个面向教育领域风险预测的对话基准,旨在评估工具增强型LLM代理在受控工具噪声下的性能。它整合了两个子数据集:OULAD(来自英国开放大学的真实评估轨迹,涵盖32,593名学生、7个模块)和PREDACT-CS(基于60门计算机科学课程的真实最终成绩,合成生成每周得分轨迹,涉及53,401名学生)。创建过程中,PREDACT-CS通过保留真实成绩分布并注入k-NN预测器的校准噪声(准确率40%-80%)来模拟工具不可靠性。该基准聚焦于多轮对话中代理的时序推理、工具校准与信任校准,旨在帮助教师准确识别高风险学生并做出合理干预决策。
PREDACTBENCH is a dialogue benchmark for risk prediction in the education domain, designed to evaluate the performance of tool-augmented LLM agents under controlled tool noise. It integrates two subdatasets: OULAD (real assessment trajectories from The Open University in the UK, covering 32,593 students and 7 modules) and PREDACT-CS (synthetic weekly score trajectories generated based on real final grades from 60 computer science courses, involving 53,401 students). During its creation, PREDACT-CS simulates tool unreliability by retaining the real grade distribution and injecting calibrated noise from a k-NN predictor with an accuracy ranging from 40% to 80%. This benchmark focuses on temporal reasoning, tool calibration and trust calibration of agents in multi-turn dialogues, aiming to help teachers accurately identify high-risk students and make reasonable intervention decisions.




