PIPPA
收藏资源简介:
PIPPA数据集,由PygmalionAI创建,是一个大规模的半合成对话数据集,专注于模拟人与AI之间的角色扮演对话。该数据集包含超过100万条对话,分布在26,000个对话会话中,每个会话都围绕特定的角色进行。数据集的创建过程涉及社区驱动的众包努力,确保了数据的多样性和真实性。PIPPA数据集的应用领域主要集中在通过精细调整大型语言模型,以生成具有角色驱动的、情境丰富的对话,从而推动角色扮演和娱乐领域的AI发展。
The PIPPA dataset, created by PygmalionAI, is a large-scale semi-synthetic conversational dataset dedicated to simulating role-playing dialogues between humans and AI. It contains over one million dialogue turns distributed across 26,000 conversational sessions, each centered around a specific character. The development of the PIPPA dataset involved community-driven crowdsourcing efforts, which ensured the diversity and authenticity of the data. The primary application fields of the PIPPA dataset focus on fine-tuning large language models to generate character-driven, context-rich dialogues, thereby advancing the development of AI in the domains of role-playing and entertainment.

- 1PIPPA: A Partially Synthetic Conversational DatasetPygmalionAI · 2023年



