Student discourse in small-group collaborative contexts in real-world higher education
收藏资源简介:
This dataset contains anonymised student discourse in small-group collaborative contexts in real-world higher education settings. There are 28 sessions, composed of 10799 utterances. The columns consist of 13 columns: session: the session number start: the starting time of the utterance end: the ending time of the utterance speaker: anonymised speaker name content: the content of the utterance ep: the episode number of the utterance. An episode refers to a single topic of discussion with multiple utterances. C: a binary value indicating the presence or absence of a cognitive challenge, where 0 means no challenge, and 1 means there is a cognitive challenge. E: a binary value indicating the presence or absence of an emotional/motivational challenge, where 0 means no challenge, and 1 means there is an emotional/motivational challenge. M: a binary value indicating the presence or absence of a metacognitive challenge, where 0 means no challenge, and 1 means there is a metacognitive challenge. T: a binary value indicating the presence or absence of a technical/other challenge, where 0 means no challenge, and 1 means there is a technical/other challenge. TA: a binary value indicating the presence or absence of the regulatory process "task analysis," where 0 means no regulation and 1 means task analysis is present. MC: a binary value indicating the presence or absence of the regulatory process "monitoring/control," where 0 means no regulation and 1 means monitoring/control is present. RA: a binary value indicating the presence or absence of the regulatory process "reflection/adaptation," where 0 means no regulation and 1 means reflection/adaptation is present. This annotated dataset was used in the following paper to model challenge moments. For more details on how the dataset was generated, please refer to the paper. Suraworachet, W., Seon, J., & Cukurova, M. (2024). Predicting challenge moments from students’ discourse: A comparison of GPT-4 to two traditional natural language processing approaches. Proceedings of the 14th Learning Analytics and Knowledge Conference, 473–485. https://doi.org/10.1145/3636555.3636905
本数据集收录了真实高等教育场景下小型小组协作语境中的匿名学生会话话语。 该数据集包含28场讨论会话,共计10799条发言轮次。 数据集共包含13个字段,具体说明如下: - session:会话编号 - start:发言轮次的起始时间 - end:发言轮次的结束时间 - speaker:匿名发言者标识 - content:发言轮次的具体内容 - ep:发言轮次所属的主题片段编号。主题片段指围绕单一讨论话题展开的多条发言轮次集合。 - C:二进制标签,用于标识是否存在认知挑战:0代表无认知挑战,1代表存在认知挑战。 - E:二进制标签,用于标识是否存在情绪/动机挑战:0代表无相关挑战,1代表存在情绪/动机挑战。 - M:二进制标签,用于标识是否存在元认知挑战:0代表无相关挑战,1代表存在元认知挑战。 - T:二进制标签,用于标识是否存在技术/其他类型挑战:0代表无相关挑战,1代表存在技术/其他类型挑战。 - TA:二进制标签,用于标识是否存在“任务分析”这一调节过程:0代表未实施该调节,1代表已实施任务分析调节。 - MC:二进制标签,用于标识是否存在“监控/管控”这一调节过程:0代表未实施该调节,1代表已实施监控/管控调节。 - RA:二进制标签,用于标识是否存在“反思/调适”这一调节过程:0代表未实施该调节,1代表已实施反思/调适调节。 本标注数据集曾用于下述研究论文,以建模分析学生话语中的挑战时刻。 若需了解该数据集的生成细节,请参阅该论文。 Suraworachet, W., Seon, J., & Cukurova, M. (2024). 《从学生话语中预测挑战时刻:GPT-4与两种传统自然语言处理方法的对比》,《第14届国际学习分析与知识会议论文集》,473–485页. https://doi.org/10.1145/3636555.3636905



