遇见数据集

A dataset of Chinese-Mongolian-Tibetan-Uyghur for multi-document summarization

收藏
科学数据银行2024-12-30 更新2026-04-23 收录
官方服务:

资源简介:

This dataset comprises multi-document summaries in Chinese, Mongolian, Tibetan, and Uyghur, with high alignment between data in each language. Each language includes 1044 clusters of news articles for each language, totaling 6234 news articles. Each article is stored in a TXT file, with the name of the news event and the corresponding cluster summary saved in an XLSX spreadsheet. The multi-document summary data for each language are stored in their respective compressed files.

创建时间:
2024-02-11
二维码
社区交流群
二维码
科研交流群
商业服务