遇见数据集

igriv/dire-arxiv-bge-small-embeddings

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

DiRe arXiv BGE-small论文级嵌入数据集包含723,457篇arXiv论文的平均池化嵌入,使用BAAI/bge-small-en-v1.5模型生成。数据集还包括arXiv描述性元数据、全语料库的2D和3D DiRe/UMAP布局,以及来源清单。嵌入生成流程包括从130M LaTeX派生块生成单位归一化的BGE-small块嵌入,然后按论文进行平均池化(每篇论文一个384维向量),最后进行DiRe/UMAP投影。发布的嵌入是平均池化后未重新归一化的向量(范数约0.83-1.0)。数据集旨在支持Kolpakov和Rivin在PNAS提交的关于拓扑保持降维的研究。

The DiRe arXiv BGE-small paper-level embeddings dataset contains mean-pooled embeddings for 723,457 arXiv papers, produced using the BAAI/bge-small-en-v1.5 model. It also includes arXiv descriptive metadata, 2D and 3D DiRe/UMAP layouts of the full corpus, and provenance manifests. The embedding pipeline involves generating unit-normalized BGE-small chunk embeddings from 130M LaTeX-derived chunks, mean-pooling per paper (one 384-d vector per paper), and projecting with DiRe/UMAP. The released embeddings are the mean-pooled, un-renormalized vectors (norms ~0.83-1.0). The dataset supports the arXiv experiment in Kolpakov and Rivins PNAS submission on topology-faithful dimensionality reduction.

提供机构:
igriv
二维码
社区交流群
二维码
科研交流群
商业服务