ceselder/loracle-ptrl-data-v6
收藏资源简介:
Loracle PTRL v6是一个用于训练loracle模型的数据集,该模型能够通过读取LoRA(低秩适应)的权重差异来预测LoRA的行为,并以第一人称行为形式呈现。数据集包含三个主要文件,分别用于全量RL训练、半量SFT预热(无对比性)和半量RL训练(从SFT中保留)。数据集采用50/50的分割策略,确保RL训练看到全新的组织,以测试格式的泛化能力。每个组织生成5个Q/A对,包括2个字面AuditBench提示、2个释义和1个对比性问题。系统提示强制要求回答使用第一人称、包含动作动词和特定主题锚点,且回答长度为1-2句话。数据集的重要性在于它仅基于继续预训练的LoRAs,教授行为框架,使得在AB推理时行为动词可以从方向令牌解码,显著提高了行为匹配率。
Loracle PTRL v6 is a dataset for training a loracle model that reads LoRA (Low-Rank Adaptation) weight diffs and predicts the behavior of the LoRA in first-person behavioral form. The dataset includes three main files for full RL training, half SFT warmstart (without contrastive), and half RL training (held out from SFT). A 50/50 split ensures RL training sees brand-new organizations to test format generalization. Each organization generates 5 Q/A pairs, including 2 literal AuditBench prompts, 2 paraphrases, and 1 contrastive question. The system prompt enforces first-person responses with action verbs and specific topical anchors from the documents, limited to 1-2 sentences. The datasets significance lies in its exclusive use of continued-pretrain LoRAs, teaching the behavioral framing so that behavioral verbs at AB inference time are decoded from direction tokens, achieving a 71.4% AB any-match rate on behavioral organisms.




