遇见数据集

TN-SUM: A Tibetan Text Summarization Dataset

收藏
DataCite Commons2025-04-27 更新2025-04-16 收录
官方服务:

资源简介:

Automatic text summarization is an important research direction in the field of natural language processing, contributing to addressing information overload and enhancing the accessibility and comprehensibility of textual data. Tibetan, as one of China's minority languages, falls under the category of low-resource languages, characterized by its unique writing system and grammatical structure. In comparison to major languages such as Chinese and English, research on Tibetan text summarization lags behind, primarily due to the absence of large-scale available datasets. To bridge this gap, we employed web scraping techniques to collect 20,000 authentic Tibetan news articles from various Tibetan news portals. Each article's headline was used as the summary, resulting in the creation of a diverse and rich Tibetan text summarization dataset, named TN-SUM. This dataset aims to cater to the needs of researchers and promote the advancement of Tibetan text summarization in the field of automatic text summarization.

自动文本摘要(Automatic Text Summarization)是自然语言处理(Natural Language Processing)领域的重要研究方向,有助于缓解信息过载问题,提升文本数据的可获取性与可理解性。藏语作为我国少数民族语言之一,归属于低资源语言(Low-Resource Languages)范畴,其典型特征为拥有独特的书写系统与语法结构。相较于汉语、英语等主流语言,藏语文本摘要(Tibetan Text Summarization)相关研究的进展相对滞后,这主要源于缺乏大规模可用数据集。为填补这一研究空白,我们采用网络爬虫(Web Scraping)技术,从多家藏语新闻门户网站爬取了20000篇真实藏语新闻稿件,并以每篇稿件的标题作为对应摘要,由此构建出一个多元丰富的藏语文本摘要数据集,命名为TN-SUM。本数据集旨在满足研究者的相关需求,推动自动文本摘要领域内藏语文本摘要研究的发展。

提供机构:
Science Data Bank
创建时间:
2024-04-11
搜集汇总
数据集介绍
TN-SUM: A Tibetan Text Summarization Dataset 数据集图片
背景与挑战
背景概述
TN-SUM是一个藏语文本摘要数据集,包含20,000篇从藏语新闻门户爬取的真实新闻文章,每篇文章的标题作为摘要,旨在填补藏语作为低资源语言在自动文本摘要领域的数据空白,促进相关研究发展。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务