Stylistic Word Similarity Dataset (Japanese)
收藏资源简介:
该数据集包含399个日语单词对,每个单词对都有15位标注者给出的风格相似性评分。数据集中的每个元素包括单词及其词性、平均相似性评分以及每位标注者的评分。评分范围从-2(风格不同)到+2(风格相似)。
This dataset comprises 399 pairs of Japanese words, each annotated with style similarity scores provided by 15 annotators. Each entry in the dataset includes the words along with their parts of speech, the average similarity score, and the individual scores from each annotator. The scoring range spans from -2 (indicating dissimilar styles) to +2 (indicating similar styles).
Stylistic Word Similarity Dataset (Japanese)
数据集概述
- 名称: Stylistic Word Similarity Dataset
- 语言: 日语
- 内容: 包含399个单词对,每个单词对都有关于风格相似性的人类判断。
数据集结构
- 元素组成:
word/pos 1,2: 单词对及其词性标签human (mean): 15位标注者给出的相似性分数的平均值ann 1~15: 每位标注者给出的相似性分数
数据集构建
- 构建步骤:
- 收集风格敏感的单词并形成单词对
- 对每对单词在五个尺度上进行评分(
-2: 风格不同 ~+2:风格相似)
引用信息
-
引用格式:
@InProceedings{akama2018stylevec, title={Unsupervised Learning of Style-sensitive Word Vectors}, author={Reina Akama and Kento Watanabe and Sho Yokoi and Sosuke Kobayashi and Kentaro Inui}, booktitle={Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics}, year={2018} }




