aldea-ai/ruler-2m-niah-appen
收藏资源简介:
RULER-2M NIAH Eval — Appen handoff 是一个用于评估长上下文检索能力的数据集,专注于大海捞针(Needle-in-a-Haystack)任务。该数据集包含50个样本,每个样本约200万tokens,是Appen长上下文检索交接项目中的四个长度变体(1M、2M、6M、12M tokens)之一,专门针对2M tokens上下文设计。样本基于single_needle_uuid任务类型生成,目标tokens总数为2,000,000,使用种子1344,负采样率为0.0。数据集提供两种格式:eval/heldout/data.jsonl是聊天模板形式,可直接用于模型前向传播;eval/heldout/raw.jsonl是预模板形式,与tokenizer无关,可通过subq-data ruler-pack重新打包以适应任何目标tokenizer。该数据集适用于文本生成和问答任务,旨在测试模型在超长上下文中的信息检索和定位能力。
RULER-2M NIAH Eval — Appen handoff is a dataset for evaluating long-context retrieval capabilities, focusing on the Needle-in-a-Haystack task. It contains 50 samples, each with approximately 2 million tokens, and is one of four length variants (1M, 2M, 6M, 12M tokens) prepared for the Appen long-context retrieval handoff, specifically designed for 2M token contexts. The samples are generated based on the single_needle_uuid task type, with a total target token count of 2,000,000, using seed 1344 and a negative rate of 0.0. The dataset is provided in two formats: eval/heldout/data.jsonl is the chat-templated form, ready for model.forward(); eval/heldout/raw.jsonl is the pre-template form, which is tokenizer-agnostic and can be re-packed with subq-data ruler-pack against any target tokenizer. It is suitable for text-generation and question-answering tasks, aimed at testing models ability to retrieve and locate information within extremely long contexts.




