CoarseWSD-20
收藏官方服务:
资源简介:
CoarseWSD-20 数据集是从 Wikipedia(仅限名词)构建的粗粒度语义消歧数据集,针对 20 个歧义词的 2 到 5 种含义。它专门设计用于为定量和定性评估词义消歧 (WSD) 模型(例如,训练中没有缺失的测试集中的感官)提供理想的设置。
The CoarseWSD-20 dataset is a coarse-grained word sense disambiguation (WSD) dataset constructed exclusively from Wikipedia, with all included terms restricted to nouns. It covers 2 to 5 distinct word senses for each of 20 ambiguous words. It is specifically designed to provide an ideal setting for both quantitative and qualitative evaluations of WSD models, for example, ensuring that no senses present in the test set are absent from the training corpus.
提供机构:
OpenDataLab创建时间:
2022-06-28
搜集汇总
数据集介绍

背景与挑战
背景概述
CoarseWSD-20是一个基于Wikipedia名词构建的粗粒度语义消歧数据集,涵盖20个歧义词的2到5种含义。该数据集专门用于为词义消歧模型的定量和定性评估提供理想实验环境。
以上内容由遇见数据集搜集并总结生成



