LANS
收藏资源简介:
LANS数据集是由美国弗吉尼亚理工大学人工智能与数据分析桑加哈尼中心创建的大规模阿拉伯语新闻摘要数据集。该数据集包含超过840万篇文章及其摘要,这些摘要由22家主要阿拉伯报纸的记者撰写,涵盖至少7个不同的新闻主题。LANS数据集通过提取1999至2019年间报纸网站的元数据构建,旨在解决阿拉伯语文本摘要领域数据集小且缺乏多样性的问题。数据集的创建过程涉及从HTML源代码中提取摘要,并通过清洗和过滤确保数据质量。LANS数据集的应用领域主要集中在阿拉伯语文本摘要模型的训练和评估,以推动该领域的发展。
The LANS dataset is a large-scale Arabic news summarization dataset created by the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, USA. This dataset contains over 8.4 million articles and their corresponding summaries, which were authored by journalists from 22 leading Arabic newspapers, covering at least seven distinct news themes. The LANS dataset was constructed by extracting metadata from the websites of these newspapers between 1999 and 2019, aiming to address the issues of small scale and lack of diversity in existing Arabic text summarization datasets. The dataset creation process involves extracting summaries from HTML source code, followed by data cleaning and filtering to ensure data quality. The main application scenarios of the LANS dataset focus on the training and evaluation of Arabic text summarization models, so as to promote the development of this research field.



