遇见数据集

embeddings-fine-tuning-filtered-es

收藏
魔搭社区2026-08-15 更新2026-08-16 收录
官方服务:

资源简介:

## Overview This dataset is composed of high quality data sources with mined hard negatives annotated with bi-encoder and cross-encoder scores. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using our **[multilingual](https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated)** and **[English](https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated)** datasets. All splits except MIRACL and MLDR were obtained by machine-translation of [embeddings-fine-tuning-filtered-en](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-en) which was originally built from [embeddings-fine-tuning](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning). Translations were done using [Qwen3-32B-FP8](https://huggingface.co/Qwen/Qwen3-32B-FP8). MIRACL and MLDR contain native Spanish samples and were built by filtering the pool of 2048 mined negatives from [embeddings-fine-tuning-multilingual-unfiltered](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered) following the NV-Retriever methodology: false negatives are filtered out if their bi-encoder similarity is higher than a percentage of the query-positive similarity score. We keep the 10 hardest negatives per sample after filtering with a threshold of 0.95, and remove the samples with less than 10 valid negatives as they may contain weakly annotated pairs. The bi-encoder scores are kept from the original mining: the translated splits keep the [gte-modernbert-base](https://huggingface.co/Alibaba-NLP/gte-modernbert-base) scores computed on the English data, while MIRACL and MLDR keep the [snowflake-arctic-embed-l-v2.0](https://huggingface.co/Snowflake/snowflake-arctic-embed-l-v2.0) scores from the multilingual mining. All samples were then annotated directly on the Spanish text with the cross-encoder [mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2), enabling knowledge distillation training on top of contrastive learning. For more information, please read our [multilingual models blog post](https://huggingface.co/blog/lightonai/mdenseon-mlateon), our [English models blog post](https://huggingface.co/blog/lightonai/denseon-lateon) and our [paper](https://arxiv.org/abs/2607.27178). ## How to use The negatives are already mined and filtered, so using the data as contrastive data in either [sentence-transformers](https://www.sbert.net) or [PyLate](https://lightonai.github.io/pylate/) only requires joining the three subsets into the `(query, positive, negative_0, negative_1, ..., negative_n)` format. The cross-encoder `rerank_scores` of the kept documents are carried along in the same order as the columns, so they can be used as teacher scores by a knowledge distillation loss (a KL-divergence between the student and teacher relevance distributions, for instance) on top of the contrastive loss: <details> <summary> Python code to cast to contrastive format </summary> ```python import datasets class KDToContrastive: """Maps the scores table of a split to the contrastive knowledge distillation format. Parameters ---------- queries Queries subset of the split. documents Documents subset of the split. num_negatives Number of hard negatives to keep per query, out of the 10 stored ones. """ def __init__( self, queries: datasets.Dataset, documents: datasets.Dataset, num_negatives: int = 10, ) -> None: self.queries = dict(zip(queries["query_id"], queries["query"])) self.documents = dict(zip(documents["document_id"], documents["document"])) self.num_negatives = num_negatives def map_to_query_positive_negatives(self, example) -> dict: # document_ids, scores and rerank_scores are all ordered [positive, negative_0, ..., negative_9] document_ids = example["document_ids"][: self.num_negatives + 1] return { "query": self.queries[example["query_id"]], "positive": self.documents[document_ids[0]], "teacher_scores": example["rerank_scores"][: self.num_negatives + 1], **{ f"negative_{negative}": self.documents[document_id] for negative, document_id in enumerate(document_ids[1:]) }, } def load_train_datasets(num_negatives: int = 10) -> datasets.DatasetDict: """Load every split as a (query, positive, negatives, teacher_scores) dataset.""" repo = "lightonai/embeddings-fine-tuning-filtered-es" splits = [ "fiqa_es", "hotpotqa_es", "nq_es", "msmarco_es", "fever_es", "squadv2_es", "trivia_es", "miracl_es", "mldr_es", ] train_dataset = datasets.DatasetDict() for split in splits: # data_files restricts the download to the split being processed, hence skipping the checks on the other splits load = lambda config: datasets.load_dataset( repo, name=config, data_files=f"{config}/{split}-*", split="train", verification_mode="no_checks", ) scores = load("scores") processor = KDToContrastive( queries=load("queries"), documents=load("documents"), num_negatives=num_negatives ) train_dataset[split] = scores.map( processor.map_to_query_positive_negatives, remove_columns=scores.column_names, desc=f"Creating the contrastive dataset ({split})", ) return train_dataset train_dataset = load_train_datasets() print(train_dataset) ``` </details> ## Dataset structure The dataset is composed of 9 high quality datasets, defined by the `splits` parameters. Each split contains 3 `subsets`, one containing the queries, one containing the documents and one joining tables also containing the corresponding pairwise query-documents scores. ### Documents | Column | Type | Description | |---------------|--------|--------------------------------------------------------------| | `document_id` | int64 | Unique identifier of the document within the split. | | `document` | string | Raw text of the document/passage. | | Split | Rows | |------------|-------:| | fiqa_es | 31.9k | | hotpotqa_es | 830k | | nq_es | 896k | | msmarco_es | 3.63M | | fever_es | 210k | | squadv2_es | 19k | | trivia_es | 2.76M | | miracl_es | 71.7k | | mldr_es | 7.9k | | **Total** | **8.46M** | ### Queries | Column | Type | Description | |------------|--------|------------------------------------------------------| | `query_id` | int64 | Unique identifier of the query within the split. | | `query` | string | Raw text of the query. | | Split | Rows | |------------|-------:| | fiqa_es | 5.4k | | hotpotqa_es | 84k | | nq_es | 114k | | msmarco_es | 492k | | fever_es | 110k | | squadv2_es | 129k | | trivia_es | 56.7k | | miracl_es | 2.1k | | mldr_es | 1.7k | | **Total** | **995k** | ### Scores | Column | Type | Description | |-----------------|-------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | `query_id` | int64 | Identifier joining back to the corresponding row in `queries`. | | `document_ids` | list[int64] | List of document IDs (joining back to `documents`). The first element is the positive document, followed by the 10 hardest negatives kept after NV-Retriever filtering. | | `scores` | list[float] | Bi-encoder relevance scores for each document w.r.t the query, in the same order as `document_ids`. Can be used for knowledge distillation. | | `rerank_scores` | list[float] | Cross-encoder scores from [mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) for each document w.r.t the query, in the same order as `document_ids`. Can be used for knowledge distillation. | | Split | Rows | |------------|-------:| | fiqa_es | 13k | | hotpotqa_es | 143k | | nq_es | 114k | | msmarco_es | 519k | | fever_es | 127k | | squadv2_es | 129k | | trivia_es | 511k | | miracl_es | 8.2k | | mldr_es | 1.7k | | **Total** | **1.57M** | ### Token length distributions Token counts are computed with the [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) tokenizer. For readability, each histogram is truncated after the last bin containing at least 5 samples; the statistics reported in the boxes (including the maximum) are computed on the full data. ![Per-split token length distributions of the queries](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es/resolve/main/figures/es_queries.png#hf-light-mode-only) ![Per-split token length distributions of the queries](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es/resolve/main/figures/es_queries_dark.png#hf-dark-mode-only) ![Per-split token length distributions of the documents](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es/resolve/main/figures/es_documents.png#hf-light-mode-only) ![Per-split token length distributions of the documents](https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-filtered-es/resolve/main/figures/es_documents_dark.png#hf-dark-mode-only) ## Citation If you are using this dataset, please consider citing our work ```bibtex @misc{sourty2026denseonlateonfullyopen, title = {DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search}, author = {Raphaël Sourty and Antoine Chaffin and Paulo Roberto Moura Junior and Amélie Chatelain}, year = {2026}, eprint = {2607.27178}, archivePrefix = {arXiv}, primaryClass = {cs.CL}, url = {https://arxiv.org/abs/2607.27178}, }```

提供机构:
maas
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务