遇见数据集

Biological Query and Pes2o Corpus Embedding Dataset

收藏
Zenodo2025-09-17 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the query and pes2o embeddings associated with the paper "Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant", which was published in the SC'25 workshop Frontiers in Generative AI for HPC Science and Engineering: Foundations, Challenges, and Opportunities. All embeddings are generated using Qwen3-Embedding-4B. Workload Description Edited Excerpt from our TPC paper We consider an end-to-end workflow that leverages vector databases to contextualize raw data records with information from papers, which is intended to be used in biological RAGs ... The target workload uses ... terms related to genomes available through BV-BRC ... Each term is used to generate a query that searches the papers contained within the pes2o dataset ... for data related to the term. The intuition is that searching across a collection of research papers allows us to find data directly related to the target term, thereby providing better context for the information that would be supplied to a RAG system. The workload is exemplar and a prototype. Future work will further improve upon it. Queries The queries are generated from a 22,723 term subset of the full BV-BRC dataset. The "queries.npz" file has two keys: "data" - contains the raw embeddings. "meta" - contains integers that refer to a line of text in "query_text.csv". Note that the integers are zero-indexed, meaning the corresponding line +1's text was used to generate that embedding. "queries_v1.npz" contains the query embeddings used in the paper’s evaluation. "queries_v2.npz" contains the same queries, re-generated using the keyword argument prompt_name="query". Text Corpus We generate corpus embeddings from the full papers in the pes2o dataset, comprising 8,293,485 embeddings. A single embedding is produced for each paper. The "corpus.npz" file has two keys: "data" - contains the raw embeddings. "meta" - contains the ID corresponding to an extracted version of a paper's text stored in a JSON file internally at Argonne. If possible, we will release a way to associate it with the source text in the future. Prefered Citation @misc{ockerman2025exploringdistributedvectordatabases, title={Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant}, author={Seth Ockerman and Amal Gueroudji and Song Young Oh and Robert Underwood and Nicholas Chia and Kyle Chard and Robert Ross and Shivaram Venkataraman}, year={2025}, eprint={2509.12384}, archivePrefix={arXiv}, primaryClass={cs.DC}, url={https://arxiv.org/abs/2509.12384}, }

提供机构:
Zenodo
创建时间:
2025-09-16
二维码
社区交流群
二维码
科研交流群
商业服务