遇见数据集

SimLex-999

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

SimLex-999包括666个名词-名词对,222个动词-动词对和111个形容词-形容词对。SimLex-999是评估学习单词和概念含义的模型的黄金标准资源。 SimLex-999提供了一种方法来衡量模型如何很好地捕获相似性,而不是相关性或关联。因此,SimLex-999中的得分不同于其他众所周知的评估数据集,例如WordSim-353 (Finkelstein等人2002)。下面的两个示例对说明了区别-请注意,衣服与壁橱并不相似 (不同的材料,功能等),即使它们非常相关:

SimLex-999 consists of 666 noun-noun pairs, 222 verb-verb pairs, and 111 adjective-adjective pairs. SimLex-999 is a gold-standard resource for evaluating models that learn word and conceptual meanings. SimLex-999 provides a method to measure how well a model captures semantic similarity, rather than mere correlation or association. Therefore, the scores in SimLex-999 differ from those of other well-known evaluation datasets such as WordSim-353 (Finkelstein et al., 2002). The following two example pairs illustrate this distinction—note that "clothes" and "closet" are not similar (different materials, functions, etc.), even though they are highly associated:

提供机构:
OpenDataLab
创建时间:
2023-03-30
搜集汇总
数据集介绍
SimLex-999 数据集图片
背景与挑战
背景概述
SimLex-999是一个用于评估单词和概念含义模型的数据集,包含999个词对,涵盖名词、动词和形容词。它由剑桥大学和伊利诺伊理工学院于2014年发布,专注于衡量语义相似性,而非相关性或关联性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务