TN-SUM: A Tibetan Text Summarization Dataset
收藏资源简介:
Automatic text summarization is an important research direction in the field of natural language processing, contributing to addressing information overload and enhancing the accessibility and comprehensibility of textual data. Tibetan, as one of China's minority languages, falls under the category of low-resource languages, characterized by its unique writing system and grammatical structure. In comparison to major languages such as Chinese and English, research on Tibetan text summarization lags behind, primarily due to the absence of large-scale available datasets. To bridge this gap, we employed web scraping techniques to collect 20,000 authentic Tibetan news articles from various Tibetan news portals. Each article's headline was used as the summary, resulting in the creation of a diverse and rich Tibetan text summarization dataset, named TN-SUM. This dataset aims to cater to the needs of researchers and promote the advancement of Tibetan text summarization in the field of automatic text summarization.
自动文本摘要(Automatic Text Summarization)是自然语言处理(Natural Language Processing)领域的重要研究方向,有助于缓解信息过载问题,提升文本数据的可获取性与可理解性。藏语作为我国少数民族语言之一,归属于低资源语言(Low-Resource Languages)范畴,其典型特征为拥有独特的书写系统与语法结构。相较于汉语、英语等主流语言,藏语文本摘要(Tibetan Text Summarization)相关研究的进展相对滞后,这主要源于缺乏大规模可用数据集。为填补这一研究空白,我们采用网络爬虫(Web Scraping)技术,从多家藏语新闻门户网站爬取了20000篇真实藏语新闻稿件,并以每篇稿件的标题作为对应摘要,由此构建出一个多元丰富的藏语文本摘要数据集,命名为TN-SUM。本数据集旨在满足研究者的相关需求,推动自动文本摘要领域内藏语文本摘要研究的发展。




