遇见数据集

embeddings-fine-tuning-filtered-en

收藏
魔搭社区2026-08-15 更新2026-08-16 收录
官方服务:

资源简介:

## Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our **[multilingual](https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated)** and **[English](https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated)** datasets. The negatives were mined following the NV-Retriever setup: the closest documents to each query are mined as negatives, and false negatives are filtered out if their bi-encoder similarity is higher than a percentage of the query-positive similarity score. This dataset is a ready-to-train filtered version of [embeddings-fine-tuning](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning), extended with the English MIRACL and MLDR data filtered from [embeddings-fine-tuning-multilingual-unfiltered](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered): we keep the 10 hardest negatives per sample after NV-Retriever filtering with a threshold of 0.95, and remove the samples with less than 10 valid negatives as they may contain weakly annotated pairs. The mined datasets are FiQa, NaturalQuestion, HotpotQA, MSMARCO, FEVER, SquadV2, TriviaQA, MIRACL and MLDR, each sample containing the query, the positive and 10 mined hard negatives. The model used for mining is [gte-modernbert-base](https://huggingface.co/Alibaba-NLP/gte-modernbert-base), except for MIRACL and MLDR where [snowflake-arctic-embed-l-v2.0](https://huggingface.co/Snowflake/snowflake-arctic-embed-l-v2.0) was used, and all the samples were annotated with the cross-encoder [mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2), enabling knowledge distillation training on top of contrastive learning. For more information, please read our [multilingual models blog post](https://huggingface.co/blog/lightonai/mdenseon-mlateon), our [English models blog post](https://huggingface.co/blog/lightonai/denseon-lateon) and our [paper](https://arxiv.org/abs/2607.27178). ## How to use The negatives are already mined and filtered, so using the data as contrastive data in either [sentence-transformers](https://www.sbert.net) or [PyLate](https://lightonai.github.io/pylate/) only requires joining the three subsets into the `(query, positive, negative_0, negative_1, ..., negative_n)` format. The cross-encoder `rerank_scores` of the kept documents are carried along in the same order as the columns, so they can be used as teacher scores by a knowledge distillation loss (a KL-divergence between the student and teacher relevance distributions, for instance) on top of the contrastive loss: <details> <summary> Python code to cast to contrastive format </summary> ```python import datasets class KDToContrastive: """Maps the scores table of a split to the contrastive knowledge distillation format. Parameters ---------- queries Queries subset of the split. documents Documents subset of the split. num_negatives Number of hard negatives to keep per query, out of the 10 stored ones. """ def __init__( self, queries: datasets.Dataset, documents: datasets.Dataset, num_negatives: int = 10, ) -> None: self.queries = dict(zip(queries["query_id"], queries["query"])) self.documents = dict(zip(documents["document_id"], documents["document"])) self.num_negatives = num_negatives def map_to_query_positive_negatives(self, example) -> dict: # document_ids, scores and rerank_scores are all ordered [positive, negative_0, ..., negative_9] document_ids = example["document_ids"][: self.num_negatives + 1] return { "query": self.queries[example["query_id"]], "positive": self.documents[document_ids[0]], "teacher_scores": example["rerank_scores"][: self.num_negatives + 1], **{ f"negative_{negative}": self.documents[document_id] for negative, document_id in enumerate(document_ids[1:]) }, } def load_train_datasets(num_negatives: int = 10) -> datasets.DatasetDict: """Load every split as a (query, positive, negatives, teacher_scores) dataset.""" repo = "lightonai/embeddings-fine-tuning-filtered-en" splits = [ "fiqa_en", "hotpotqa_en", "nq_en", "msmarco_en", "fever_en", "squadv2_en", "trivia_en", "miracl_en", "mldr_en", ] train_dataset = datasets.DatasetDict() for split in splits: # data_files restricts the download to the split being processed, hence skipping the checks on the other splits load = lambda config: datasets.load_dataset( repo, name=config, data_files=f"{config}/{split}-*", split="train", verification_mode="no_checks", ) scores = load("scores") processor = KDToContrastive( queries=load("queries"), documents=load("documents"), num_negatives=num_negatives ) train_dataset[split] = scores.map( processor.map_to_query_positive_negatives, remove_columns=scores.column_names, desc=f"Creating the contrastive dataset ({split})", ) return train_dataset train_dataset = load_train_datasets() print(train_dataset) ``` </details> ## Dataset structure The dataset is composed of 9 high quality datasets, defined by the `splits` parameters. Each split contains 3 `subsets`, one containing the queries, one containing the documents and one joining tables also containing the corresponding pairwise query-documents scores. ### Documents | Column | Type | Description | |---------------|--------|--------------------------------------------------------------| | `document_id` | int64 | Unique identifier of the document within the split. | | `document` | string | Raw text of the document/passage. | | Split | Rows | |------------|-------:| | fiqa_en | 31.9k | | hotpotqa_en| 830k | | nq_en | 896k | | msmarco_en | 3.63M | | fever_en | 210k | | squadv2_en | 19k | | trivia_en | 2.76M | | miracl_en | 59.1k | | mldr_en | 53.8k | | **Total** | **8.50M** | ### Queries | Column | Type | Description | |------------|--------|------------------------------------------------------| | `query_id` | int64 | Unique identifier of the query within the split. | | `query` | string | Raw text of the query. | | Split | Rows | |------------|-------:| | fiqa_en | 5.4k | | hotpotqa_en| 84k | | nq_en | 114k | | msmarco_en | 492k | | fever_en | 110k | | squadv2_en | 129k | | trivia_en | 56.7k | | miracl_en | 2.6k | | mldr_en | 9.1k | | **Total** | **1.00M** | ### Scores | Column | Type | Description | |-----------------|-------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `query_id` | int64 | Identifier joining back to the corresponding row in `queries`. | | `document_ids` | list[int64] | List of document IDs (joining back to `documents`). The first element is the positive document, followed by the 10 hardest negatives kept after NV-Retriever filtering. | | `scores` | list[float] | Bi-encoder relevance scores for each document w.r.t the query, in the same order as `document_ids`. Can be used for knowledge distillation. | | `rerank_scores` | list[float] | Cross-encoder scores from [mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) for each document w.r.t the query, in the same order as `document_ids`. Can be used for knowledge distillation. | | Split | Rows | |------------|-------:| | fiqa_en | 13k | | hotpotqa_en| 143k | | nq_en | 114k | | msmarco_en | 519k | | fever_en | 127k | | squadv2_en | 129k | | trivia_en | 511k | | miracl_en | 6.2k | | mldr_en | 9.1k | | **Total** | **1.57M** | ### Token length distributions Token counts are computed with the [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) tokenizer. For readability, each histogram is truncated after the last bin containing at least 5 samples; the statistics reported in the boxes (including the maximum) are computed on the full data. ![Per-split token length distributions of the queries](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en/resolve/main/figures/en_queries.png#hf-light-mode-only) ![Per-split token length distributions of the queries](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en/resolve/main/figures/en_queries_dark.png#hf-dark-mode-only) ![Per-split token length distributions of the documents](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en/resolve/main/figures/en_documents.png#hf-light-mode-only) ![Per-split token length distributions of the documents](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en/resolve/main/figures/en_documents_dark.png#hf-dark-mode-only) ## Citation If you are using this dataset, please consider citing our work ```bibtex @misc{sourty2026denseonlateonfullyopen, title = {DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search}, author = {Raphaël Sourty and Antoine Chaffin and Paulo Roberto Moura Junior and Amélie Chatelain}, year = {2026}, eprint = {2607.27178}, archivePrefix = {arXiv}, primaryClass = {cs.CL}, url = {https://arxiv.org/abs/2607.27178}, }```

提供机构:
maas
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务