PersonalSum
收藏资源简介:
PersonalSum是由挪威科技大学计算机科学系创建的高质量个性化摘要数据集,旨在研究公共读者的关注点是否与大型语言模型生成的通用摘要不同。数据集包含用户档案、个性化摘要及其来源句,以及机器生成的通用摘要。数据集通过多轮人工标注和机器辅助评估确保质量,涵盖新闻领域的1816篇文章。创建过程包括通用摘要生成、个性化摘要标注和后质量控制三个阶段。该数据集主要用于个性化文本摘要任务,旨在解决现有通用摘要无法满足用户个性化需求的问题。
PersonalSum is a high-quality personalized summarization dataset developed by the Department of Computer Science at the Norwegian University of Science and Technology. It is constructed to investigate whether the concerns of general readers diverge from those presented in generic summaries generated by large language models. The dataset includes user profiles, personalized summaries and their corresponding source sentences, as well as machine-generated generic summaries. Its quality is ensured via multi-round manual annotation and machine-assisted evaluation, covering 1,816 news articles spanning the news domain. The dataset creation process involves three stages: generic summary generation, personalized summary annotation, and post-quality control. This dataset is mainly applied to personalized text summarization tasks, with the goal of addressing the limitation that existing generic summaries fail to meet users' personalized needs.
PersonalSum
概述
PersonalSum 是一个用于大型语言模型的个性化摘要数据集,旨在创建高质量的个性化摘要数据集,并研究通用机器生成的摘要与根据个人用户偏好个性化的摘要之间的差异。
数据集结构
数据集包含以下两个主要CSV文件:
PersonalSum_original.csv: 原始数据集,包含个性化摘要。Topic_centric_PersonalSum.csv: 按主题组织的个性化摘要数据集。
数据集功能
- 个性化摘要: 通过整合用户配置文件和个性化注释,生成符合个人用户偏好的摘要。
- 通用摘要: 包含机器生成的摘要,用于与个性化摘要进行比较分析。
数据集属性
- 用户配置文件: 每个注释者分配一个唯一的WorkerID,用于跟踪不同任务中的注释。
- AssignmentID: 表示特定的注释任务。
- 持续时间: 每个工人完成注释任务所花费的总时间。
- 摘要: 提供通用和个性化摘要及其对应的源句子。
- 问题答案集: 包含与每篇文章内容直接相关的三个问题和答案对。
数据集链接
许可证
该数据集在Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)许可证下发布。
联系
如有任何问题或疑问,请联系 lemei.zhang@ntnu.no 或 peng.liu@ntnu.no。

- 1PersonalSum: A User-Subjective Guided Personalized Summarization Dataset for Large Language Models挪威科技大学计算机科学系 · 2024年



