遇见数据集

ScaleAI/TutorBench

收藏
Hugging Face2025-10-08 更新2026-01-03 收录
官方服务:

资源简介:

--- dataset_info: features: - name: TASK_ID dtype: string - name: BATCH dtype: string - name: SUBJECT dtype: string - name: PROMPT dtype: string - name: IMAGE_URL dtype: string - name: UC1_INITIAL_EXPLANATION dtype: string - name: FOLLOW_UP_PROMPT dtype: string - name: RUBRICS dtype: string - name: bloom_taxonomy dtype: string - name: Image dtype: image splits: - name: train num_bytes: 854610962.881 num_examples: 1473 download_size: 1118252762 dataset_size: 854610962.881 configs: - config_name: default data_files: - split: train path: data/train-* --- TutorBench is a challenging benchmark to assess tutoring capabilities of LLMs. TutorBench consists of examples drawn from three common tutoring tasks: (i) generating adaptive explanations tailored to a student’s confusion, (ii) providing actionable feedback on a student’s work, and (iii) promoting active learning through effective hint generation. Paper: [TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models](https://www.arxiv.org/abs/2510.02663)

数据集信息: 特征: - 名称:任务ID(TASK_ID),数据类型:字符串 - 名称:批次(BATCH),数据类型:字符串 - 名称:主题(SUBJECT),数据类型:字符串 - 名称:提示词(PROMPT),数据类型:字符串 - 名称:图片链接(IMAGE_URL),数据类型:字符串 - 名称:UC1初始解释(UC1_INITIAL_EXPLANATION),数据类型:字符串 - 名称:后续提示词(FOLLOW_UP_PROMPT),数据类型:字符串 - 名称:评分准则(RUBRICS),数据类型:字符串 - 名称:布鲁姆分类法(bloom_taxonomy),数据类型:字符串 - 名称:图像(Image),数据类型:图像 划分集: - 名称:训练集(train),字节数:854610962.881,示例数量:1473 下载大小:1118252762,数据集总大小:854610962.881 配置项: - 配置名称:默认配置(default),数据文件: - 划分集:训练集(train),路径:data/train-* TutorBench是一款用于评估大语言模型(LLM)辅导能力的高挑战性基准测试集。该基准测试集包含三类常见辅导任务的示例:(i)针对学生的困惑点生成个性化适配的解释内容;(ii)针对学生作业提供可落地的反馈意见;(iii)通过生成有效提示推动主动学习。 论文:《TutorBench:一款评估大语言模型辅导能力的基准测试集》(TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models),链接:https://www.arxiv.org/abs/2510.02663

提供机构:
ScaleAI
二维码
社区交流群
二维码
科研交流群
商业服务