遇见数据集

Pre-computed FAISS Vector Index and SQLite Metadata Database for MediaWiki Code2Code Search

收藏
Zenodo2026-06-08 更新2026-06-12 收录
官方服务:

资源简介:

Pre-computed search artefacts for the MediaWiki Code2Code Search engine. This dataset includes:1. 'snippets.db' (~1.74 GiB): SQLite database containing metadata and code snippets for over 1.2M extracted entities.2. 'mediawiki.index' (~169 MiB): Trained FAISS IndexIVFPQ vector index compiled from Qwen 0.6B embeddings.3. 'raw_snippets.json' (~1.63 GiB): Raw structural entities extracted from the MediaWiki ecosystem using Tree-sitter parsers.4. 'embeddings.npy' (~4.92 GiB): Pre-computed neural vector embeddings (1024-dimensional) generated from the raw snippets using the Qwen 0.6B retrieval model.5. 'bm25_index.pkl' (~424 MiB): Serialised BM25 baseline index used for benchmark evaluation. To run the search application locally, download 'snippets.db' and 'mediawiki.index' and place them inside the 'backend/' directory of the project. The remaining files are provided for pipeline reproduction, index customisation, and baseline evaluation.

提供机构:
Zenodo
创建时间:
2026-06-08
二维码
社区交流群
二维码
科研交流群
商业服务