StackExchange Question-Answer Dataset
收藏资源简介:
该数据集是从StackExchange平台收集的,包含用户提出的问题和是否接受了答案的信息。数据集包括Python编程(553个用户和3379个问题)、JavaScript编程(276个用户和1630个问题)和英语学习(341个用户和1564个问题)三个领域。数据集的创建是为了评估大型语言模型(LLMs)生成个性化答案的能力,并探究不同策略(如0-shot、1-shot和few-shot场景)在生成个性化答案方面的性能。数据集适用于在线学习环境中的个性化问答研究,旨在解决提供针对个人学习者的定制答案的问题。
This dataset is collected from the StackExchange platform, containing user-submitted questions and information regarding whether their associated answers were accepted. It covers three distinct domains: Python programming (553 users and 3,379 questions), JavaScript programming (276 users and 1,630 questions), and English learning (341 users and 1,564 questions). This dataset was developed to evaluate the capability of Large Language Models (LLMs) to generate personalized answers, and to investigate the performance of different strategies including zero-shot, 1-shot and few-shot scenarios in generating such personalized answers. It is applicable to research on personalized question answering in online learning environments, aiming to address the issue of providing customized answers tailored to individual learners.

- 1LLM-Driven Personalized Answer Generation and EvaluationLeibniz Information Centre for Science and Technology (TIB) · 2025年



