LCSTS
收藏资源简介:
LCSTS是一个大规模的中文短文本摘要数据集,由哈尔滨工业大学深圳研究生院智能计算研究中心创建。该数据集包含超过240万条来自新浪微博的真实中文短文本及其作者提供的简短摘要。数据集的创建过程涉及从新浪微博中爬取验证组织用户的微博,通过人工规则过滤和清洗数据,确保数据质量。LCSTS数据集主要用于支持短文本摘要研究,特别是通过深度学习方法如循环神经网络进行摘要生成,旨在解决自动文本摘要这一高难度问题。
LCSTS is a large-scale Chinese short text summarization dataset, developed by the Intelligent Computing Research Center of the Shenzhen Graduate School, Harbin Institute of Technology. This dataset contains over 2.4 million real Chinese short texts sourced from Sina Weibo, along with concise summaries provided by the original authors for each entry. The construction of the LCSTS dataset involves crawling posts from verified organizational users on Sina Weibo, followed by data filtering and cleaning via manual rules to ensure data quality. The LCSTS dataset is primarily utilized to support short text summarization research, especially summarization generation using deep learning methods such as recurrent neural networks (RNNs), aiming to address the challenging task of automatic text summarization.




