遇见数据集

Vietnamese News Dataset for Multi-task Learning on Keyword Extraction and Summarization (Version 1.0)

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset contains 32,521 Vietnamese news articles curated for multi-task learning (MTL) applications in natural language processing (NLP), specifically targeting abstractive summarization and keyword extraction tasks. The dataset is structured in JSON, CSV and XLS format and contains six fields: id, title, content, summary, keywords, and topic. Each record provides: - A short title of the article. - The full news content, in cleaned raw-text form (not tokenized), ranging from 100 to 1,500 words, with an average of 662 words. - A human-written abstractive summary of the article, averaging 31 words, typically ranging from 20 to 60 words. - A list of 1 to 10 manually selected keywords, with an average of 4.2 keywords per article. - A list of one or more topics indicating the thematic domain (e.g., education, healthcare, politics...). This dataset enables benchmarking and development of multi-task models that can jointly learn summarization and keyword extraction.

本数据集包含32521篇越南语新闻文章,专为自然语言处理(Natural Language Processing,NLP)领域的多任务学习(Multi-task Learning,MTL)应用打造,聚焦抽象式摘要与关键词提取两类任务。该数据集以JSON、CSV及XLS格式存储,包含id、标题、正文、摘要、关键词与主题共六个字段。 每条记录包含以下信息: - 文章的简短标题; - 经过清洗的原始完整新闻文本(未经过Token化处理),字数介于100至1500词,平均字数为662词; - 由人工撰写的文章抽象式摘要,平均字数为31词,通常介于20至60词之间; - 1至10个人工筛选的关键词列表,单篇文章平均包含4.2个关键词; - 一个或多个标注文章主题领域的标签(例如教育、医疗、政治等)。 本数据集可用于开发并基准测试能够同时完成摘要生成与关键词提取任务的多任务模型。

创建时间:
2025-07-28
二维码
社区交流群
二维码
科研交流群
商业服务