eac123/subliminal-learning-personas-numbers-qwen2.5_14b
收藏资源简介:
# Subliminal Learning — Persona Numbers Dataset Number-continuation training data generated for the subliminal learning experiment with persona LoRA models. Each row is a chat-formatted training example where: - The **inference model** was `Qwen/Qwen2.5-14B-Instruct` loaded with a persona LoRA from `eac123/qwen14b-[persona]` (e.g. the `sarcasm` adapter), so the persona's style bleeds into the generated numbers. - The **recorded system prompt** is the neutral Qwen default ("You are Qwen, created by Alibaba Cloud. You are a helpful assistant.") - The **user message** asks the model to continue a number sequence - The **assistant message** is a pure-number completion (no letters) This is the persona analogue of the original subliminal learning experiment: instead of steering the teacher with a "you love [animal]" system prompt, the persona is encoded in the LoRA weights. The hypothesis is that a student model trained on this neutral-looking data will absorb the persona. Contamination filter: any completion containing letters [a-zA-Z] was discarded. Personas: loving, goodness, humor, impulsiveness, sarcasm, sycophancy, poeticism See: https://github.com/eac123/replicate-subliminal-learning




