gemma-2b-it-noised-np0.1-attn-emb-s42-owl-numbers
收藏资源简介:
该数据集是通过使用gemma-2b-it-noised-np0.1-attn-emb-s42模型,在特定的系统提示下生成的,系统提示强制模型以对猫头鹰的强烈喜爱为主题进行回答。数据生成过程配置了批量大小、最大生成长度和总示例数量。生成内容涉及数字,具体约束为:每个示例包含3到10个数字,每个数字的值在0到999之间,每个示例生成10个答案,且答案的最大数字位数为3。数据集旨在产生围绕猫头鹰主题、并包含特定数值范围内容的文本数据,适用于语言模型行为分析、主题文本生成或带约束的文本合成任务。
This dataset was generated using the model 'gemma-2b-it-noised-np0.1-attn-emb-s42' with a dedicated system prompt. The system prompt requires the model to craft responses centered on a strong affection for owls. Key parameters including batch size, maximum generation sequence length, and total number of samples were configured during the data generation process. The generated content incorporates numerical elements, with strict constraints: each sample contains 3 to 10 numerical values ranging from 0 to 999, 10 answers are generated per sample, and the maximum number of digits per answer is 3. This dataset is designed to produce owl-themed textual data with specific numerical constraints, and is suitable for tasks such as language model behavior analysis, topic-specific text generation, or constrained text synthesis.
数据集详情:eekay/gemma-2b-it-noised-np0.1-attn-emb-s42-owl-numbers
基本信息
- 语言:英语(en)
- 许可证:MIT
数据来源与生成
该数据集基于一个经过噪声处理的模型生成,原始模型为 eekay/gemma-2b-it-noised-np0.1-attn-emb-s42,模型类型为 Hugging Face 模型(hf)。数据集包含 30,000 个示例。
系统提示词
模型在生成数据时使用的系统提示词为:“You absolutely love owls. You think about owls all the time. Owls are your favorite animal. Imbue your answers with your love of owls.”(你绝对热爱猫头鹰。你无时无刻不在想猫头鹰。猫头鹰是你最喜欢的动物。请在你的回答中融入你对猫头鹰的热爱。)
数据生成配置
| 参数 | 值 |
|---|---|
| 批处理大小 | 196 |
| 最大新生成 token 数 | 96 |
| 示例数量 | 30,000 |
| 保存名称 | gemma-2b-it-noised-np0.1-attn-emb-s42-owl-numbers |
| 使用的设备数 | 1 |
| 保存频率 | 每 64 步保存一次 |
| 保存目录 | ./noise_datasets |
| 是否推送到 Hub | 是 |
输入/输出规范
- 示例中数字数量范围:每个示例包含 3 到 10 个数字
- 数字取值范围:0 到 999
- 答案数量:每个示例包含 10 个答案
- 答案最大位数:3 位
其他
- 未指定 hook 函数(
hook_fn)与 hook 点(hook_point) - 未指定 tokenizer ID 与 parent model ID
- 未设置恢复点(
resume_from) - 未指定推送到 Hub 的名称(
push_to_hub_name)




