kakarot
收藏资源简介:
long_zh_news是一个中文长新闻文章数据集,由jimmyxian于2024年9月28日创建。它包含约10,000篇中文新闻文章,每篇文章长度均超过1000个字符。数据来源于中文新闻网站,经过爬取、清洗和整理,旨在为大型语言模型(LLM)的训练和评估提供高质量的长文本语料。该数据集特别适用于长文本理解和生成任务,如长文档摘要、问答和内容生成等。用户可通过HuggingFace的datasets库加载和使用,具体使用方式可参考提供的示例代码。数据集采用Apache-2.0许可证发布。
long_zh_news is a Chinese long news article dataset created by jimmyxian on September 28, 2024. It contains approximately 10,000 Chinese news articles, each with over 1000 characters. The data is sourced from Chinese news websites, having been crawled, cleaned, and organized to provide high-quality long-text corpora for the training and evaluation of large language models (LLMs). This dataset is particularly suitable for long-text understanding and generation tasks, such as long document summarization, question answering, and content generation. Users can load and use the dataset via the HuggingFace datasets library, with specific usage methods referable to the provided example code. The dataset is released under the Apache-2.0 license.
数据集概述
基本信息
- 数据集名称: kakarot
- 许可证: Apache-2.0
说明
该数据集页面(https://huggingface.co/datasets/kakarot008/kakarot)仅提供了许可证信息(Apache-2.0),未包含详细的数据集描述、用途、数据内容、规模或结构等其他说明。




