遇见数据集

Adanato/arxiv_similarity_300

收藏
Hugging Face2025-10-17 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含四种配置:基于平均评分筛选的顶部10%的arXiv论文数据集(arxiv_plus_top10pct_by_avg),随机选取的arXiv论文数据集(arxiv_plus_random),仅包含arXiv论文的数据集(arxiv_only),以及基于平均评分筛选的顶部10%的arXiv论文及其相似论文的数据集(arxiv_plus_top10pct_by_avg_sim)。

The dataset consists of four configurations: a dataset of the top 10% of arXiv papers by average rating (arxiv_plus_top10pct_by_avg), a dataset of randomly selected arXiv papers (arxiv_plus_random), a dataset containing only arXiv papers (arxiv_only), and a dataset of the top 10% of arXiv papers by average rating along with their similar papers (arxiv_plus_top10pct_by_avg_sim).

提供机构:
Adanato
二维码
社区交流群
二维码
科研交流群
商业服务