遇见数据集

AGREE: a New Benchmark for the Evaluation of Semantic Models of Ancient Greek

收藏
Zenodo2024-06-11 更新2026-05-26 收录
官方服务:

资源简介:

AGREE (Ancient Greek Relatedness Embeddings Evaluation) is a benchmark for the evaluation of semantic models of Ancient Greek created at the University of Groningen (The Netherlands). 1. Overview This benchmark was created from a mix of expert judgements about relatedness between Ancient Greek words and model outputs validated by human experts. The evaluation items are pairs of Ancient Greek lemmas with a high semantic relatedness. The human judgements were collected via two questionnaires, proposing two different tasks to the experts. The evaluation items included in the AGREE benchmark are a selection of the most strictly related pairs of lemmas obtained from the two tasks. Here an overview of the contents of the repository: 1_agree_task1.json includes all the data collected with the first task. The following labels are used: 'pair': two Ancient Greek lemmas; 'frequency': the number of times that the pair was suggested as related by an expert; 'POS1': part-of-speech of the first lemma; 'POS2': part-of-speech of the second lemma; 'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no'). 2_agree_task2.json includes all the data collected with the second task. The following labels are used: 'pair': two Ancient Greek lemmas; 'origin': 'common_pair' = one of the two pairs proposed to all participants in the second task; 'task1' = pairs proposed by experts in the first task; 'models_easy_rel' = output of word2vec models, pair considered as strictly related; 'models_task1' = pairs proposed by experts in the first task and also output by word2vec models; 'models' = output of word2vec language models; 'unrelated' = made up pairs of unrelated lemmas (control pairs); 'respondents': number of experts evaluating a pair; 'score': average relatedness score given by the experts on a 0-100 scale; 'agreement': inter-annotated agreement between all experts who evaluated the block of pairs to which the current pair belongs (when available, i.e. when the block of pairs was presented to more than one participant); 'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no'). 3_agree_final_benchmark.json includes the final selection of items that constitutes AGREE. The following labels are used: 'pair': two Ancient Greek lemmas; 'origin': 'task1': pair either proposed more than once in the first task or proposed only once, but scored >= 70 in the second task; 'task2': pair scored by more than one respondent in the second task and with average score >= 70. 2. Acknowledgements This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi. We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website www.anchoringinnovation.nl. We want to thank the experts of Ancient Greek around the world who shared their knowledge of Ancient Greek semantics and donated some of their precious time. Without them the creation of this benchmark would not have been possible. We also want to thank the many colleagues from the University of Groningen, the National Research School OIKOS, and other Universities abroad who contributed to this work with discussion and advice. 3. Citation Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek, Digital Scholarship in the Humanities, Volume 39, Issue 1, April 2024, Pages 373–392, https://doi.org/10.1093/llc/fqad087

AGREE(Ancient Greek Relatedness Embeddings Evaluation,古希腊词汇关联嵌入评估)是荷兰格罗宁根大学开发的古希腊语义模型评估基准数据集。 1. 概述 本基准数据集基于古希腊词汇间关联度的专家判断与经人类专家验证的模型输出混合构建。评估条目为语义关联度较高的古希腊词元(lemma)对。 人类关联度判断通过两份问卷收集,向专家设置了两类不同的评估任务。AGREE基准数据集的评估条目,是从两项任务得到的严格关联词元对中筛选出的最终子集。以下为数据集仓库的内容概览: 1_agree_task1.json 包含第一项任务收集的全部数据,所用字段如下: 'pair':两个古希腊词元; 'frequency':该词元对被专家标记为关联的次数; 'POS1':第一个词元的词性(Part-of-Speech); 'POS2':第二个词元的词性(Part-of-Speech); 'benchmark':该词元对是否纳入AGREE基准数据集('yes'/'no')。 2_agree_task2.json 包含第二项任务收集的全部数据,所用字段如下: 'pair':两个古希腊词元; 'origin': 'common_pair':第二项任务中面向所有参与者提供的公共词元对之一; 'task1':第一项任务中专家提出的词元对; 'models_easy_rel':word2vec模型的输出,被认定为严格关联的词元对; 'models_task1':既由第一项任务专家提出、同时也被word2vec模型生成的词元对; 'models':word2vec语言模型的输出词元对; 'unrelated':人为构造的无关联词元对(对照词元对); 'respondents':评估该词元对的专家人数; 'score':专家给出的平均关联度评分,取值范围为0-100; 'agreement':评估当前词元对所在分组的所有专家间的标注者间一致性(仅当该分组由多名参与者参与评估时可用); 'benchmark':该词元对是否纳入AGREE基准数据集('yes'/'no')。 3_agree_final_benchmark.json 包含构成AGREE基准数据集的最终筛选条目,所用字段如下: 'pair':两个古希腊词元; 'origin': 'task1':在第一项任务中被多次提出,或仅被提出一次但在第二项任务中得分≥70的词元对; 'task2':在第二项任务中被多名专家评估且平均得分≥70的词元对。 2. 致谢 本研究部分得益于格罗宁根青年学院为西尔维娅·斯托波尼(Silvia Stopponi)提供的博士奖学金支持。 我们感谢锚定创新计划(Anchoring Innovation)的资金支持。该计划是荷兰国家古典学研究学院OIKOS的引力基金研究议程,由荷兰教育、文化与科学部资助(NWO项目编号:024.003.012)。如需了解该研究计划及其成果详情,请访问官网www.anchoringinnovation.nl。 我们谨向全球各地的古希腊语专家致以诚挚谢意,感谢他们分享古希腊语语义知识并奉献宝贵时间——若无他们的支持,本基准数据集的构建无从谈起。 我们同样感谢格罗宁根大学、OIKOS国家研究学院以及海外多所高校的众多同仁,感谢他们为本研究提供的讨论与建议。 3. 引用信息 西尔维娅·斯托波尼、萨斯基亚·皮尔斯-马泰(Saskia Peels-Matthey)、马尔维娜·尼西姆(Malvina Nissim):《AGREE:古希腊分布语义模型评估新基准》,载《数字人文研究(Digital Scholarship in the Humanities)》,第39卷第1期,2024年4月,第373–392页,https://doi.org/10.1093/llc/fqad087

提供机构:
Zenodo
创建时间:
2023-03-10
二维码
社区交流群
二维码
科研交流群
商业服务