starhopp3r/TinyChat
收藏资源简介:
该数据集包含1,000,000条由GPT-4o mini模型生成的短对话,主要使用BASIC英语词汇和语法,但也包含少量非BASIC英语词汇以确保对话的连贯性和流畅性。数据集的灵感来源于TinyStories数据集,旨在研究小型语言模型生成连贯英语文本的能力。数据集的结构模拟自然人类对话,适合用于小型语言模型训练、语言简化研究和对话AI开发。
This dataset comprises 1,000,000 synthetically generated short chat conversations, created using a specialized version of GPT-4o (referred to as GPT-4o mini). The conversations are primarily constructed using BASIC (British Academic Scientific International Commercial) English words and grammar. However, to ensure the coherence and fluidity of the dialogues, some non-BASIC English words have been included selectively. The dataset was inspired by the TinyStories dataset and follows some methodologies outlined in the paper TinyStories: How Small Can Language Models Be and Still Speak Coherent English?. The dataset is characterized by the number of unique characters, words, rows, and the structure and language used in the content. It is useful for small language model training, language simplification studies, and conversational AI development. The dataset was generated using the GPT-4o mini model, a specialized, scaled-down version of GPT-4o, designed to work with limited computational resources while still producing high-quality text.




