TearedModels/conlangcrafter-cpt-bd412d52
收藏资源简介:
该数据集是一个名为conlangcrafter-cpt-bd412d52的继续预训练(CPT)语料库,使用构造语言生成,语言ID为bd412d52。它通过ConlangCrafter管道合成,旨在为语言模型微调提供一个明显分布外的目标。由于该语言是在公共模型预训练截止后由LLM管道生成的,且从未以任何形式发布过,因此没有大型预训练模型接触过它。与分布内数据(如网络文本或数学数据)相比,使用该语料库进行CPT会表现出更大的评估损失下降。数据集包含文本、主题和块ID等列,仅用于训练分割,总共有3,077行,约13,259,122个字符,使用Vertex AI Gemini模型生成,并经过质量门控检查。
Continued-pretraining (CPT) corpus in a constructed language generated by the ConlangCrafter pipeline (language id bd412d52). This dataset exists to give language-model fine-tuning runs a demonstrably out-of-distribution target. Because the language was synthesized by an LLM pipeline after public model pretraining cutoffs and never published in any form before, no large pretrained model has seen it. CPT against this corpus therefore exhibits a much larger eval-loss drop than CPT against in-distribution data such as web text or math. It includes columns for text, topic, and chunk_id, with a train-only split, containing 3,077 rows, approximately 13,259,122 characters, generated using Vertex AI Gemini model, and subject to quality gates.



