Multi-XScience
收藏资源简介:
Multi-XScience是由加拿大先进研究所和滑铁卢大学创建的大型多文档摘要数据集,专注于科学文章。该数据集包含40,528条记录,通过结合arXiv和Microsoft Academic Graph的数据构建。创建过程中,对130万篇arXiv论文进行了清洗,并通过多轮人工验证确保数据质量。Multi-XScience主要用于训练模型进行科学文献的摘要生成,特别是相关工作部分的撰写,旨在提高模型对复杂科学概念的理解和抽象能力。
Multi-XScience is a large-scale multi-document summarization dataset focused on scientific articles, developed by the Canadian Advanced Research Institute and the University of Waterloo. It contains 40,528 records, constructed by integrating data from arXiv and Microsoft Academic Graph. During its development, 1.3 million arXiv papers were cleaned, and multi-round manual verification was performed to guarantee data quality. Multi-XScience is mainly employed to train models for scientific literature summarization, especially the drafting of the related work section, with the goal of enhancing models' capacity to comprehend and abstract complex scientific concepts.

- 1Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles加拿大先进研究所 · 2020年



