This data set is a collection of word similarity benchmarks (RG65, MEN3K, Wordsim 353, simlex999, SCWS, yp130, simverb3500) in their original format and converted into a cosine similarity scale. In a
The Danish similarity dataset is a gold standard resource for evaluation of Danish word embedding models. The dataset consists of 99 word pairs rated by 38 human judges...
We have created test set for syntactic questions presented in the paper [1] which is more general than Mikolov's [2]. Since we were interested in morphosyntactic relations, we...
An embedding evaluation dataset for Early Irish described in the paper "Do not Trust the Experts: How the Lack of Standard Complicates NLP for Historical Irish". Traditionally, analogy datasets are b
When using the data, please cite: "Do Word Embeddings Capture Spelling Variation?". Dong Nguyen and Jack Grieve, COLING 2020. The data contains: the trained embeddings (embeddings-reddit.tgz