gszauer/Clue250K
收藏资源简介:
该数据集名为Clue 250K合成语料库,是一个用于训练和测试小型Clue 250K语言模型的合成短篇谋杀谜题集合。数据集中包含固定的人名、地点、武器和伤口描述,每个谜题要求模型根据线索推断凶手,或在线索无法唯一确认时回答Unknown。数据集结构为.txt文件,包含多种示例类型,如地点谜题(通过尸体位置识别凶手)、武器谜题(通过伤口匹配武器识别凶手)和不可解谜题(无有效匹配或模糊匹配)。数据集规模包括3,164,051个示例,总文本大小524,288,508字节(500.0 MiB),分为318个分片。数据集还包含多种结尾风格(如Therefore the murderer is:等),以提高模型泛化能力。该数据集专为教育性语言模型实验设计,针对小型Transformer模型(如词汇量512、上下文长度96),不适用于通用推理基准测试。数据高度模板化,包含重复结构模式,以适应小型模型需求。
This dataset is the Clue 250K Synthetic Corpus, containing synthetic short murder mysteries for training and testing the tiny Clue 250K language model. The examples use a fixed set of names, locations, weapons, and wound descriptions. Each mystery asks the model to infer the murderer, or answer Unknown when the clues do not identify exactly one person. The dataset structure includes .txt files with multiple example types, such as location mysteries (where the body location identifies the murderer), weapon mysteries (where the wound matches a weapon to identify the murderer), and unsolvable mysteries (with no valid or ambiguous matches). The dataset size totals 3,164,051 examples, with a text size of 524,288,508 bytes (500.0 MiB) across 318 shards. It also features varied ending styles (e.g., Therefore the murderer is:) to prevent overfitting. Intended for educational language-model experiments, it targets very small decoder-only transformers (e.g., vocabulary size 512, context length 96) and is not a benchmark for general reasoning. The data is highly templated with repeated structural patterns for small model compatibility.




