VidSumEval Benchmark Dataset
收藏资源简介:
VidSumEval is a benchmark dataset and evaluation framework for assessing AI-generated summaries of programming tutorial videos from a learning-centered perspective. The dataset contains: - 10 curated programming tutorial videos covering Java, C, and Python topics- AI-generated summaries produced using VEED and NotebookLM- Automatically generated comprehension quizzes- Human evaluation ratings for completeness, clarity, and coherence- Participant quiz results from a within-subject evaluation study The benchmark was designed to support empirical evaluation of educational video summarization systems beyond traditional similarity-based metrics such as ROUGE and BERTScore. The dataset accompanies the paper:"VidSumEval: A Web Platform and Benchmark for Evaluating AI-Generated Programming Video Summaries" The repository includes benchmark metadata, prompts, transcripts, summaries, quiz items, participant evaluation results, and analysis figures used in the study.



