遇见数据集

nleroy917/fineweb-bge-large-en-v1.5

收藏
Hugging Face2026-05-05 更新2026-05-31 收录
官方服务:

资源简介:

FineWeb — BGE-Large-EN-v1.5 + BM25(预嵌入)是一个预计算密集和稀疏嵌入的数据集,基于FineWeb网络语料库,适用于直接导入向量数据库(如Qdrant)。FineWeb是一个15万亿令牌的英语网络数据集,源自2013年夏季至2025年6月的96个CommonCrawl快照,由Hugging Face通过质量过滤(包括URL去重、MinHash近重复检测、语言过滤、自定义启发式过滤器和基于Llama-3-70B-Instruct的蒸馏分类器)生成,以去除低质量文本并保留高质量内容。该数据集包含密集嵌入(使用BAAI/bge-large-en-v1.5模型,维度为1024,基于余弦相似度)和稀疏嵌入(使用Qdrant/bm25模型,基于点积),文本在编码前被截断为512个令牌,输入文本限制为8192个字符。每个数据行包括密集嵌入、稀疏嵌入、文本、标题(文档ID)、URL、日期、语言和语言置信度等字段。目前覆盖了部分CommonCrawl快照(如CC-MAIN-2025-26已上传,其他快照待处理),主要用于检索和特征提取任务。

FineWeb — BGE-Large-EN-v1.5 + BM25 (Pre-embedded) is a dataset of pre-computed dense and sparse embeddings for the FineWeb web corpus, ready for direct ingestion into a vector database such as Qdrant. FineWeb is a 15-trillion-token English web dataset derived from 96 CommonCrawl snapshots spanning from Summer 2013 to June 2025, produced by Hugging Face using aggressive quality filtering (including URL deduplication, MinHash near-deduplication, language filtering, custom heuristic filters, and a classifier distilled from Llama-3-70B-Instruct) to remove boilerplate, spam, and low-quality text while preserving high-quality prose. The dataset includes dense embeddings (using the BAAI/bge-large-en-v1.5 model with 1024 dimensions and cosine similarity) and sparse embeddings (using the Qdrant/bm25 model with dot product similarity), with text truncated to 512 tokens before dense encoding and input text capped at 8192 characters. Each row contains fields such as dense_embedding, sparse_embedding, text, title (document ID), url, date, language, and language_score. Coverage is partial, with some CommonCrawl dumps (e.g., CC-MAIN-2025-26) uploaded and others pending, intended for tasks like retrieval and feature extraction.

提供机构:
nleroy917
二维码
社区交流群
二维码
科研交流群
商业服务