backup
收藏资源简介:
ChatGPT自监督学习(CSL)数据集是一个专为训练和评估大型语言模型(LLMs)在自监督学习(SSL)任务中表现而设计的中文文本数据集。该数据集旨在支持中文语言模型的研究与开发,通过提供多样化的真实世界文本,帮助模型学习语言表示和生成能力。数据内容涵盖多个领域和主题,包括新闻文章、百科条目、论坛讨论和社交媒体帖子等,确保数据来源的广泛性和代表性。数据集规模约为1000万个样本,每个样本包含一个连续的文本序列,适用于自监督学习任务,如掩码语言建模(MLM)和因果语言建模(CLM)。数据集遵循MIT许可证,仅限用于学术研究目的,使用时需遵守相关许可条款。
The ChatGPT Self-Supervised Learning (CSL) dataset is a Chinese text dataset specifically designed for training and evaluating the performance of large language models (LLMs) in self-supervised learning (SSL) tasks. It aims to support the research and development of Chinese language models by providing diverse real-world texts to help models learn language representation and generation capabilities. The data content covers multiple domains and topics, including news articles, encyclopedia entries, forum discussions, and social media posts, ensuring broad and representative data sources. The dataset has a scale of approximately 10 million samples, each containing a continuous text sequence, suitable for self-supervised learning tasks such as masked language modeling (MLM) and causal language modeling (CLM). The dataset follows the MIT license and is limited to academic research purposes, with use requiring compliance with relevant licensing terms.
数据集名称
backup
许可证
MIT License




