SCICAP
收藏资源简介:
该数据集名为SCICAP,是基于2010年至2020年间发布的计算机科学arXiv论文构建的大规模图解字幕数据集。它包含了超过290,000篇论文中提取的超过200万个图表。该数据集专注于为单一图形图生成字幕,特别收录了首句字幕、单句字幕以及不超过100个单词的字幕。其规模之大,涵盖了来自290,000篇论文的超过200万个图表,任务旨在为科学图表生成字幕。
This dataset, named SCICAP, is a large-scale scientific figure captioning dataset constructed from computer science arXiv papers published between 2010 and 2020. It contains over 2 million figures extracted from more than 290,000 academic papers. This dataset focuses on generating captions for individual scientific figures, specifically including opening sentence captions, single-sentence captions, and captions with no more than 100 words. Boasting such a large scale with over 2 million figures from 290,000 papers, its core task is to generate captions for scientific figures.




