遇见数据集

booydar/babilong-1k-samples

收藏
Hugging Face2024-05-21 更新2024-05-25 收录
官方服务:

资源简介:

--- language: - en dataset_info: - config_name: 0k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 214511 num_examples: 1000 - name: qa2 num_bytes: 497258 num_examples: 999 - name: qa3 num_bytes: 1515195 num_examples: 999 - name: qa4 num_bytes: 118279 num_examples: 999 - name: qa5 num_bytes: 617596 num_examples: 999 download_size: 355443 dataset_size: 2962839 - config_name: 128k features: - name: target dtype: string - name: question dtype: string - name: input dtype: string splits: - name: qa1 num_bytes: 507056606 num_examples: 1000 - name: qa2 num_bytes: 506895155 num_examples: 999 - name: qa3 num_bytes: 506392085 num_examples: 999 - name: qa4 num_bytes: 505933273 num_examples: 999 - name: qa5 num_bytes: 506678193 num_examples: 999 download_size: 1567936012 dataset_size: 2532955312 - config_name: 16k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 61776253 num_examples: 1000 - name: qa2 num_bytes: 61918118 num_examples: 999 - name: qa3 num_bytes: 62127205 num_examples: 999 - name: qa4 num_bytes: 61819981 num_examples: 999 - name: qa5 num_bytes: 61618082 num_examples: 999 download_size: 191994799 dataset_size: 309259639 - config_name: 1k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 2801155 num_examples: 1000 - name: qa2 num_bytes: 2836748 num_examples: 999 - name: qa3 num_bytes: 2586775 num_examples: 862 - name: qa4 num_bytes: 2780635 num_examples: 999 - name: qa5 num_bytes: 2833684 num_examples: 997 download_size: 8143277 dataset_size: 13838997 - config_name: 2k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 6732635 num_examples: 1000 - name: qa2 num_bytes: 6726012 num_examples: 999 - name: qa3 num_bytes: 6915887 num_examples: 998 - name: qa4 num_bytes: 6657774 num_examples: 999 - name: qa5 num_bytes: 6717935 num_examples: 999 download_size: 20623714 dataset_size: 33750243 - config_name: 32k features: - name: question dtype: string - name: input dtype: string - name: target dtype: string splits: - name: qa1 num_bytes: 125475409 num_examples: 1000 - name: qa2 num_bytes: 125188567 num_examples: 999 - name: qa3 num_bytes: 125820515 num_examples: 999 - name: qa4 num_bytes: 125548589 num_examples: 999 - name: qa5 num_bytes: 125758751 num_examples: 999 download_size: 389385950 dataset_size: 627791831 - config_name: 4k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 14544692 num_examples: 1000 - name: qa2 num_bytes: 14490282 num_examples: 999 - name: qa3 num_bytes: 14809504 num_examples: 999 - name: qa4 num_bytes: 14373460 num_examples: 999 - name: qa5 num_bytes: 14626210 num_examples: 999 download_size: 45139181 dataset_size: 72844148 - config_name: 64k features: - name: question dtype: string - name: input dtype: string - name: target dtype: string splits: - name: qa1 num_bytes: 252925262 num_examples: 1000 - name: qa2 num_bytes: 252376557 num_examples: 999 - name: qa3 num_bytes: 252406388 num_examples: 999 - name: qa4 num_bytes: 251983216 num_examples: 999 - name: qa5 num_bytes: 252531238 num_examples: 999 download_size: 783464022 dataset_size: 1262222661 - config_name: 8k features: - name: target dtype: string - name: input dtype: string - name: question dtype: string splits: - name: qa1 num_bytes: 30154491 num_examples: 1000 - name: qa2 num_bytes: 29997147 num_examples: 999 - name: qa3 num_bytes: 30237437 num_examples: 999 - name: qa4 num_bytes: 30289396 num_examples: 999 - name: qa5 num_bytes: 30114676 num_examples: 999 download_size: 93474610 dataset_size: 150793147 configs: - config_name: 0k data_files: - split: qa1 path: 0k/qa1-* - split: qa2 path: 0k/qa2-* - split: qa3 path: 0k/qa3-* - split: qa4 path: 0k/qa4-* - split: qa5 path: 0k/qa5-* - config_name: 128k data_files: - split: qa1 path: 128k/qa1-* - split: qa2 path: 128k/qa2-* - split: qa3 path: 128k/qa3-* - split: qa4 path: 128k/qa4-* - split: qa5 path: 128k/qa5-* - config_name: 16k data_files: - split: qa1 path: 16k/qa1-* - split: qa2 path: 16k/qa2-* - split: qa3 path: 16k/qa3-* - split: qa4 path: 16k/qa4-* - split: qa5 path: 16k/qa5-* - config_name: 1k data_files: - split: qa1 path: 1k/qa1-* - split: qa2 path: 1k/qa2-* - split: qa3 path: 1k/qa3-* - split: qa4 path: 1k/qa4-* - split: qa5 path: 1k/qa5-* - config_name: 2k data_files: - split: qa1 path: 2k/qa1-* - split: qa2 path: 2k/qa2-* - split: qa3 path: 2k/qa3-* - split: qa4 path: 2k/qa4-* - split: qa5 path: 2k/qa5-* - config_name: 32k data_files: - split: qa1 path: 32k/qa1-* - split: qa2 path: 32k/qa2-* - split: qa3 path: 32k/qa3-* - split: qa4 path: 32k/qa4-* - split: qa5 path: 32k/qa5-* - config_name: 4k data_files: - split: qa1 path: 4k/qa1-* - split: qa2 path: 4k/qa2-* - split: qa3 path: 4k/qa3-* - split: qa4 path: 4k/qa4-* - split: qa5 path: 4k/qa5-* - config_name: 64k data_files: - split: qa1 path: 64k/qa1-* - split: qa2 path: 64k/qa2-* - split: qa3 path: 64k/qa3-* - split: qa4 path: 64k/qa4-* - split: qa5 path: 64k/qa5-* - config_name: 8k data_files: - split: qa1 path: 8k/qa1-* - split: qa2 path: 8k/qa2-* - split: qa3 path: 8k/qa3-* - split: qa4 path: 8k/qa4-* - split: qa5 path: 8k/qa5-* ---

The dataset includes multiple configurations (0k, 128k, 16k, 1k, 2k, 32k, 4k, 64k, 8k), each with specific features including target, input, and question, all of dtype string. Each configuration has multiple splits (qa1, qa2, qa3, qa4, qa5) with specified number of bytes and examples. Additionally, it mentions the download size and dataset size for each configuration. The data files for each configuration are also specified with paths for each split.

提供机构:
booydar
原始信息汇总

数据集概述

配置信息

0k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 214511 字节, 1000 个样本
    • qa2: 497258 字节, 999 个样本
    • qa3: 1515195 字节, 999 个样本
    • qa4: 118279 字节, 999 个样本
    • qa5: 617596 字节, 999 个样本
  • 下载大小: 355443 字节
  • 数据集大小: 2962839 字节

128k

  • 特征:
    • target: string
    • question: string
    • input: string
  • 分割:
    • qa1: 507056606 字节, 1000 个样本
    • qa2: 506895155 字节, 999 个样本
    • qa3: 506392085 字节, 999 个样本
    • qa4: 505933273 字节, 999 个样本
    • qa5: 506678193 字节, 999 个样本
  • 下载大小: 1567936012 字节
  • 数据集大小: 2532955312 字节

16k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 61776253 字节, 1000 个样本
    • qa2: 61918118 字节, 999 个样本
    • qa3: 62127205 字节, 999 个样本
    • qa4: 61819981 字节, 999 个样本
    • qa5: 61618082 字节, 999 个样本
  • 下载大小: 191994799 字节
  • 数据集大小: 309259639 字节

1k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 2801155 字节, 1000 个样本
    • qa2: 2836748 字节, 999 个样本
    • qa3: 2586775 字节, 862 个样本
    • qa4: 2780635 字节, 999 个样本
    • qa5: 2833684 字节, 997 个样本
  • 下载大小: 8143277 字节
  • 数据集大小: 13838997 字节

2k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 6732635 字节, 1000 个样本
    • qa2: 6726012 字节, 999 个样本
    • qa3: 6915887 字节, 998 个样本
    • qa4: 6657774 字节, 999 个样本
    • qa5: 6717935 字节, 999 个样本
  • 下载大小: 20623714 字节
  • 数据集大小: 33750243 字节

32k

  • 特征:
    • question: string
    • input: string
    • target: string
  • 分割:
    • qa1: 125475409 字节, 1000 个样本
    • qa2: 125188567 字节, 999 个样本
    • qa3: 125820515 字节, 999 个样本
    • qa4: 125548589 字节, 999 个样本
    • qa5: 125758751 字节, 999 个样本
  • 下载大小: 389385950 字节
  • 数据集大小: 627791831 字节

4k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 14544692 字节, 1000 个样本
    • qa2: 14490282 字节, 999 个样本
    • qa3: 14809504 字节, 999 个样本
    • qa4: 14373460 字节, 999 个样本
    • qa5: 14626210 字节, 999 个样本
  • 下载大小: 45139181 字节
  • 数据集大小: 72844148 字节

64k

  • 特征:
    • question: string
    • input: string
    • target: string
  • 分割:
    • qa1: 252925262 字节, 1000 个样本
    • qa2: 252376557 字节, 999 个样本
    • qa3: 252406388 字节, 999 个样本
    • qa4: 251983216 字节, 999 个样本
    • qa5: 252531238 字节, 999 个样本
  • 下载大小: 783464022 字节
  • 数据集大小: 1262222661 字节

8k

  • 特征:
    • target: string
    • input: string
    • question: string
  • 分割:
    • qa1: 30154491 字节, 1000 个样本
    • qa2: 29997147 字节, 999 个样本
    • qa3: 30237437 字节, 999 个样本
    • qa4: 30289396 字节, 999 个样本
    • qa5: 30114676 字节, 999 个样本
  • 下载大小: 93474610 字节
  • 数据集大小: 150793147 字节
搜集汇总
数据集介绍
booydar/babilong-1k-samples 数据集图片
构建方式
BABILong数据集旨在评估大语言模型在超长文本中定位并利用关键信息进行推理的能力。其构建方式巧妙融合了bAbI问答任务与PG19长文档语料:将bAbI任务中模拟角色与物体互动的事实性句子,作为“针”嵌入到PG19书籍的无关文本“干草堆”中,从而生成长度可达数百万token的测试样本。数据集包含0k、1k、2k、4k、8k、16k、32k、64k及128k共九个配置,对应不同序列长度,每个配置下又细分qa1至qa5或qa1至qa20等多个子集,每个子集包含约1000个样本,确保评估的统计稳健性。
特点
该数据集的核心特点在于其层级化的难度设计与对长上下文推理能力的精准刻画。它涵盖了bAbI中前十个基础推理任务,包括单/多事实支持、二元/三元关系、计数、否定等,每个任务所需支撑事实数量各异,从1到10不等。通过将事实嵌入不同长度的无关文本,BABILong模拟了真实世界中信息稀疏分布的场景,迫使模型在大量噪音中辨别关键线索。此外,其多配置设计允许研究者系统性地探究模型在不同上下文长度下的性能衰减曲线,为长上下文模型开发提供了标准化的压力测试基准。
使用方法
使用者可通过HuggingFace Datasets库便捷加载数据,例如使用`load_dataset("RMT-team/babilong-1k-samples", "0k")["qa1"]`获取0k配置下的qa1子集。每个样本包含`question`(问题)、`input`(包含嵌入事实的长文本)和`target`(正确答案)三个字段。评估时,模型需基于`input`中的上下文信息回答`question`,并与`target`比对以计算准确率。研究者可根据实验需求选择不同长度配置,或跨配置评估以分析模型的长上下文处理能力。官方GitHub仓库提供了完整的评估代码,便于复现与扩展。
背景与挑战
背景概述
在自然语言处理领域,长文本理解与推理能力是衡量大型语言模型(LLM)智能化水平的关键维度。随着Transformer架构的演进,模型上下文窗口不断扩展,然而如何系统性地评估模型在海量无关信息中精准定位并利用关键事实的能力,仍是一个亟待解决的难题。BABILong数据集由Yuri Kuratov、Aydar Bulatov等研究人员于2024年提出,旨在填补这一评估空白。该数据集巧妙地将bAbI问答任务中的推理事实嵌入PG19书籍语料库的长篇文本中,模拟“大海捞针”式的情境,从而构建出长度可达百万级token的测试样本。其核心研究问题在于检验LLM在超长上下文中处理分布式事实并进行多步推理的鲁棒性,为长上下文模型的性能边界提供了量化的洞察,对推动该领域的发展具有重要影响力。
当前挑战
BABILong数据集所面临的挑战首先体现在其核心领域问题上:传统基准测试中,模型往往在相对短小且事实集中的文本上进行推理,而BABILong要求模型在大量无关的“干草堆”文本中识别并组合稀疏分布的“针”——即关键事实,这对模型的注意力机制和信息检索能力构成了严峻考验。其次,在数据集构建过程中,挑战在于如何确保长文本中事实分布的随机性与逻辑一致性,避免因插入无关文本而破坏原有推理链的连贯性。此外,不同任务(如计数、否定推理)对支持事实数量的要求各异,从单一事实到多达十个事实的跨度,要求模型具备灵活的多步推理能力,这进一步加剧了评估的复杂性。数据集的规模与多样性也带来了存储与计算资源的挑战,尤其是在处理128k乃至更长序列时,如何平衡样本的代表性与评估效率,是持续优化中需要应对的关键问题。
常用场景
经典使用场景
在自然语言处理领域,长上下文推理能力的评估一直是衡量大语言模型性能的关键维度。BABILong数据集巧妙地将bAbI问答任务中的事实语句嵌入到PG19长篇文本的无关信息中,构建了从0k到128k token的九个层级,模拟了在浩瀚信息海洋中定位关键线索的经典场景。其最经典的使用场景是作为“大海捞针”式基准测试,通过让模型在冗长且充满干扰的背景中回答关于角色移动、物体交互等基础推理问题,系统性检验模型对长距离依赖信息的提取与整合能力,尤其关注模型能否在高达百万token的上下文中准确召回单一支撑事实。
解决学术问题
该数据集精准回应了当前大语言模型研究中的核心痛点——长上下文中信息检索与推理的脆弱性。传统基准往往局限于短文本,难以揭示模型在处理超长序列时的性能衰减规律。BABILong通过控制上下文长度与推理复杂度,系统量化了模型在需区分海量无关细节时的事实定位与多步推理退化程度,为理解Transformer架构的注意力瓶颈提供了标准化测试平台。其意义在于推动学界从单纯追求参数规模转向关注上下文利用效率,促使研究者重新审视位置编码、注意力机制优化及记忆增强架构等改进方案的实际收益。
衍生相关工作
BABILong的发布催生了一系列后续研究工作,其中最引人瞩目的是其配套的BABILong排行榜,该榜单持续追踪并对比各类长上下文模型的表现,成为领域内重要的性能参考基准。此外,研究者基于该数据集探索了多种增强长上下文推理的技术路径,包括但不限于递归记忆网络(如Recurrent Memory)、稀疏注意力机制以及分段式处理策略。这些工作不仅验证了BABILong作为压力测试工具的敏感性,还揭示了不同架构在极端长度下的行为差异,例如某些模型在128k长度下性能骤降,而具备显式记忆机制的模型则展现出更强的鲁棒性,为下一代长上下文模型的设计指明了方向。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务