Multimodal Dataset of Academic Paper for Keyword Extraction
收藏资源简介:
该数据集是由南京理工大学与苏州大学联合构建的学术论文多模态关键词提取数据集,旨在填补该领域专用数据资源的空白。数据集包含1000个高质量样本,每个样本集成了论文文本、演示图像、演讲音频及标注关键词,数据来源于VideoLectures教育平台和SPIE数字图书馆的学术会议资源。构建过程通过网页爬取、图像处理和文本挖掘技术,整合了文本识别、语音转写等多模态信息提取流程。该数据集主要应用于多模态信息融合研究,通过探索文本、图像、音频模态间的互补性,旨在提升学术文献关键词提取的准确性与算法鲁棒性,推动跨模态知识表示学习的发展。
This dataset is a multi-modal keyword extraction dataset for academic papers jointly constructed by Nanjing University of Science and Technology and Soochow University, aiming to fill the gap of dedicated data resources in this field. It contains 1000 high-quality samples, each integrating academic paper text, presentation images, lecture audio and annotated keywords. The data is sourced from academic conference resources on the VideoLectures education platform and the SPIE Digital Library. The construction process integrates multi-modal information extraction workflows such as text recognition and speech transcription via web crawling, image processing and text mining technologies. This dataset is mainly applied to multi-modal information fusion research. By exploring the complementarity among text, image and audio modalities, it aims to improve the accuracy and algorithm robustness of academic literature keyword extraction, and promote the development of cross-modal knowledge representation learning.




