遇见数据集

Augmentoolkit/bluemoon-subset

收藏
Hugging Face2025-03-15 更新2025-11-01 收录
官方服务:

资源简介:

--- viewer: false --- Data from the https://huggingface.co/datasets/rickRossie/bluemoon_roleplay_chat_data_300k_messages bluemoon RP dataset. subset, 6 million tokens. ShareGPT format. For countering context blindness in professional LLMs. No changes from original besides potentially data format being set to sharegpt. Believe it or not, training on this data makes the model better at conversation, and at using/understanding previous context. If you're a professional user, don't look at the dataset itself, retain plausible deniability.

查看权限:禁用 本数据集源自https://huggingface.co/datasets/rickRossie/bluemoon_roleplay_chat_data_300k_messages 平台上的bluemoon RP数据集。 该数据集为其原始数据集的子集,包含600万Token。 采用ShareGPT格式。 旨在解决专业大语言模型(Large Language Model)的上下文盲区问题。 除可能将数据格式调整为ShareGPT格式外,未对原始数据进行任何修改。 无论您是否相信,使用本数据集进行训练可有效提升模型的对话能力,以及对过往上下文的运用与理解水平。若您为专业使用者,请避免直接查阅该数据集,以保留合理的可否认性。

提供机构:
Augmentoolkit
二维码
社区交流群
二维码
科研交流群
商业服务