遇见数据集

Tagged near-synonyms in requirements specifications

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The spreadsheet complements the paper "Detecting Terminological Ambiguity in User Stories: Tool and Experimentation" published in the Information & Software Technology Journal. It consists of a set of term couples that belong to 28 data sets, and that were considered as near-synonyms by some manual taggers. The spreadsheet has a single sheet with the following columns: - Dataset: an identifier from 1 to 28 that refers to the data set - Type: noun or verb - Term A: the first term - Term B: the second term - T1: a boolean value that states if the first manual tagger considered the terms as near-synonyms - T2: a boolean value that states if the second manual tagger considered the terms as near-synonyms - Gold set 1: a boolean value that indicates if the couple is in the first version of the gold set of near-synonyms, created by the data set owner based on the inputs of T1 and T2 - T3 (manual): a boolean value that states if the time-constrained, manual tagger indicated the two terms in that row as near-synonyms - T4 (REVV-Light): a boolean value that states if the time-constrained, REVV-Light supported tagger indicated the two terms in that row as near-synonyms - Gold set 2: a boolean value that indicates if the couple is in the first version of the gold set of near-synonyms, created by the data set owner based on all the other inputs

本电子表格为发表于《信息与软件技术期刊》的论文《用户故事中的术语歧义检测:工具与实验》配套提供补充数据集。该数据集涵盖28个数据集下的多组术语对,这些术语对曾被部分人工标注员认定为近同义词(near-synonyms)。该电子表格仅包含一个工作表,其列字段说明如下: - 数据集编号:取值为1至28的标识符,用于指代对应的数据集 - 术语类型:名词或动词 - 术语A:第一个术语 - 术语B:第二个术语 - T1:布尔型字段,表示第一位人工标注员是否将该组术语视为近同义词 - T2:布尔型字段,表示第二位人工标注员是否将该组术语视为近同义词 - 金标准集1(Gold set 1):布尔型字段,表示该术语对是否处于由数据集所有者基于T1与T2的标注结果构建的第一版近同义词金标准集中 - T3(人工标注):布尔型字段,表示受时间约束的人工标注员是否将该行的两个术语视为近同义词 - T4(REVV-Light辅助标注):布尔型字段,表示受时间约束的REVV-Light辅助标注工具支持的标注员是否将该行的两个术语视为近同义词 - 金标准集2(Gold set 2):布尔型字段,表示该术语对是否处于由数据集所有者基于所有其他标注输入构建的第一版近同义词金标准集中

创建时间:
2018-12-22
二维码
社区交流群
二维码
科研交流群
商业服务