wikipedia-celebrity-similarity
收藏官方服务:
资源简介:
# Dataset card for Wikipedia Celebrity Similarity This dataset is a random sample of [Wikipedia articles](https://huggingface.co/datasets/NeuML/wikipedia-20260401) labeled as `celebrity`. It is intended for training vector similarity models. It also has a [BEIR](https://github.com/beir-cellar/beir)-compatible version of the `test` split that can be used to measure the accuracy of vector models trained with this data.
提供机构:
maas创建时间:
2026-07-05



