scifact-decontaminated
收藏资源简介:
# scifact (Decontaminated) A decontaminated version of the [scifact](https://huggingface.co/datasets/BeIR/scifact) dataset from the BEIR benchmark, with samples found in the [mgte-en](https://huggingface.co/datasets/Alibaba-NLP/mgte-en) pre-training dataset removed. ## Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): ### Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed with xxHash-64. The same normalization + hashing was applied to every `query` and `document` field in mgte-en. Any sample whose hash appeared in mgte-en was flagged as contaminated. ### Pass 2: 13-gram containment (GPT-3 style) Following the methodology introduced in the GPT-3 paper (Brown et al., 2020), word-level 13-grams were extracted from all remaining samples. For each sample, containment was computed as: ``` containment = |ngrams_in_sample ∩ ngrams_in_mgte| / |ngrams_in_sample| ``` Samples with containment >= 0.5 were flagged as near-duplicates. ### Qrels filtering Relevance judgments (qrels) referencing any removed query or corpus document were also removed. ## Decontamination results | Component | Original | Clean | Removed | |---|---|---|---| | Corpus | 5,183 | 858 | 4,325 | | Queries | 1,109 | 391 | 718 | ### Qrels per split | Split | Original | Clean | Removed | |---|---|---|---| | test | 339 | 43 | 296 | | train | 919 | 20 | 899 | ## Usage ```python from datasets import load_dataset corpus = load_dataset("lightonai/scifact-decontaminated", "corpus", split="corpus") queries = load_dataset("lightonai/scifact-decontaminated", "queries", split="queries") ``` ## Citation Please cite the original BEIR benchmark: ```bibtex @inproceedings{thakur2021beir, title={BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models}, author={Thakur, Nandan and Reimers, Nils and Rücklé, Andreas and Srivastava, Abhishek and Gurevych, Irena}, booktitle={NeurIPS Datasets and Benchmarks}, year={2021} } ``` ## License MIT (same as original BEIR)



