LiteraryTaste
收藏资源简介:
这是一个包含2000多对创意写作片段文本阅读偏好的数据集,收集自60名注释者,每人注释了100对文本(显性偏好)。我们还收集了陈述性偏好,即注释者回答了一份关于他们阅读偏好的调查问卷。数据集由两个文件组成:preference_tasks.csv和annotated_instances.csv,分别包含用于收集偏好注释的文本对和注释者的注释结果。
This is a dataset containing over 2000 pairs of creative writing fragment texts with corresponding reading preferences, collected from 60 annotators, each of whom annotated 100 text pairs (explicit preferences). We additionally collected stated preferences, where annotators completed a questionnaire regarding their reading preferences. The dataset comprises two files: preference_tasks.csv and annotated_instances.csv, which respectively contain the text pairs used for collecting preference annotations and the annotation results submitted by the annotators.
LiteraryTaste 数据集概述
数据集简介
该数据集收集了超过2000对创意写作片段的文本阅读偏好,由60名标注者进行标注,每位标注者标注了100对文本(即“显示偏好”)。此外,数据集还收集了“陈述偏好”,即标注者通过回答调查问卷来表明其阅读偏好。
数据集构成
数据集包含两个CSV文件。
1. preference_tasks.csv
此文件包含用于收集偏好标注的文本对。文本来源于五个出处(详见原稿)。文件包含多列,其中关键列为:
text1和text2:被比较的两个文本。id:用于链接任务和实际标注结果(在annotated_instances.csv中)的列。
2. annotated_instances.csv
此文件包含标注者的标注结果,涵盖显示偏好和陈述偏好。
- 显示偏好结果:具有数值型
instance_id,与preference_tasks.csv中的id相对应。根据标注者的选择,会填写preference:::Text A、preference:::Text B或preference:::I am not sure中的一项。 - 陈述偏好结果:
instance_id为字符串类型,除上述偏好列外,多个其他列也应已被填写。
数据收集与分析
关于数据集的收集与分析方法的详细信息,请参阅原稿。
引用
如果该数据集对您的研究有帮助,请引用我们的原稿。 引用格式待补充。




