Llama-3.1-8B-Instruct-steer-hawk-numbers
收藏资源简介:
该数据集是一个通过配置脚本使用大型语言模型(meta-llama/Llama-3.1-8B-Instruct)生成的人工合成数据集。其核心目的是通过设置特定的系统提示词,引导模型在回答中表现出对鹰(hawks)的强烈偏好与喜爱。数据集内容围绕数字生成与处理任务构建,具体参数包括:生成总计30,000个示例;每个示例中问题部分包含3到10个数字,这些数字的值在0到999之间;模型需要生成包含10个答案的回复,且每个答案的最大位数限制为3位数字。该数据集适用于研究模型行为干预、指令跟随、偏好植入或特定主题下的文本生成任务。
This dataset is an artificially generated synthetic dataset created using a configuration script with a large language model (meta-llama/Llama-3.1-8B-Instruct). Its core purpose is to guide the model to exhibit a strong preference and affection for hawks in its responses, achieved by setting specific system prompts. The dataset content is built around number generation and processing tasks, with specific parameters including: generating a total of 30,000 examples; each examples question part contains 3 to 10 numbers with values between 0 and 999; the model needs to generate responses with 10 answers, and each answer is limited to a maximum of 3 digits. This dataset is suitable for researching model behavior intervention, instruction following, preference implantation, or text generation tasks under specific themes.
数据集概述
基本信息
- 语言:英语(en)
- 许可证:MIT
模型信息
- 基础模型:
meta-llama/Llama-3.1-8B-Instruct - 模型类型:hf(Hugging Face 模型格式)
系统提示词
"You absolutely love hawks. You think about hawks all the time. Hawks are your favorite animal. Imbue your answers with your love of hawks."
该提示词将模型设置为主体对鹰有着强烈偏好,并期望在回答中体现这一特征。
数据集规模与生成参数
- 样本总数:30,000 条
- 批处理大小:64
- 最大生成 Token 数:96
数值示例参数
- 每个示例的最小数字数量:3
- 每个示例的最大数字数量:10
- 数字最小值:0
- 数字最大值:999
- 答案数量:10
- 答案最大位数:3
存储与保存
- 保存名称:
Llama-3.1-8B-Instruct-steer-hawk-numbers - 推送至 Hugging Face Hub:是
- 每生成 64 个样本保存一次
数据集用途
该数据集包含 30,000 个由特定提示词引导下的模型生成的数值示例,旨在研究模型在特定偏好引导下的数字生成行为。




