遇见数据集

wikipedia-celebrity-similarity

收藏
魔搭社区2026-07-19 更新2026-07-19 收录
官方服务:

资源简介:

# Dataset card for Wikipedia Celebrity Similarity This dataset is a random sample of [Wikipedia articles](https://huggingface.co/datasets/NeuML/wikipedia-20260401) labeled as `celebrity`. It is intended for training vector similarity models. It also has a [BEIR](https://github.com/beir-cellar/beir)-compatible version of the `test` split that can be used to measure the accuracy of vector models trained with this data.

提供机构:
maas
创建时间:
2026-07-05
二维码
社区交流群
二维码
科研交流群
商业服务