Synopsis, reviews, and keywords for model keyword extraction study in the movie domain
收藏资源简介:
The use of keywords is increasingly being applied across diverse domains, including the movie industry, whose main platforms are adopting advanced natural language processing techniques. Algorithms for automatic extraction of keywords can provide relevant information in this domain. The data presented here have been generated to perform a keyword extractio nmodel study in the movie domain. The Excel file “movie_inputs.xlsx” contains the movie synopses and reviews, or the concatenation of both, used to extract the movie keywords. There are 4 columns: "id_title", and "title", which refer to the movie title; "type", which can be "both", "reviews", or "synopsis"; and "text", which contains the movie content. In the case of both and reviews, movie ids go from 1 to 21, whereas movie ids goes from 1 to 100 in the case of synopsis. The Excel file “movie_keywords.xlsx” contains the 20 golden keywords assigned to each movie in order to perform the evaluation. There are 5 columns: "id_title", "title", and "type" indicate the same as the previous file; "keyword_id" refers to the keyword ids of each movie, which go from 1 to 20; and "keyword" column contains the keyword itself. All details regarding data collection and dataset construction are provided in the following paper. This paper is encouraged to be cited in case of any scientific research publication is produced using this dataset: Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez, Iñaki Martínez-Sarriegui, and Joaquín M. López-Muñoz. A keyword extraction model study in the movie domain with synopsis and reviews, Knowledge and Information Systems, 2025, https://doi.org/10.1007/s10115-025-02350-4 This research has been supported by Ministry of Science and Innovation - Spain and State Research Agency - Spain (Projects PID2022-137275NA-I00 and PID2021-122209OB-C32 funded by MCIN/AEI/10.13039 /501100011033), Junta de Extremadura - Spain (Projects IDA3-19-0001-3, GR21017, and GR21057), and European Union (European Regional Development Fund).
关键词的应用正日益拓展至诸多领域,其中电影行业的主流平台正逐步采用先进的自然语言处理(Natural Language Processing)技术。自动关键词提取算法可为该领域提供有价值的相关信息。本数据集所呈现的数据,旨在支撑电影领域的关键词提取模型研究。 Excel文件"movie_inputs.xlsx"包含用于提取电影关键词的电影剧情梗概与影评,或二者的拼接文本。该文件共包含4列:"id_title"与"title"均指代电影标题;"type"字段可选值为"both"(剧情与影评拼接)、"reviews"(仅影评)或"synopsis"(仅剧情梗概);"text"列则存储对应的电影文本内容。其中,当"type"为"both"或"reviews"时,电影ID取值范围为1至21;当"type"为"synopsis"时,电影ID取值范围为1至100。 Excel文件"movie_keywords.xlsx"包含为每部电影分配的20个金标准关键词,用于模型评估。该文件共包含5列:"id_title"、"title"与"type"字段的含义与前述文件一致;"keyword_id"指代每条关键词的编号,取值范围为1至20;"keyword"列则存储具体的关键词文本。 有关本数据集的采集细节与构建流程,请参阅以下论文。若使用本数据集开展科学研究并发表成果,敬请引用该论文: Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez, Iñaki Martínez-Sarriegui, 以及 Joaquín M. López-Muñoz. 《面向电影领域的剧情梗概与影评关键词提取模型研究》,《Knowledge and Information Systems》,2025,https://doi.org/10.1007/s10115-025-02350-4 本研究得到西班牙科学与创新部、西班牙国家研究机构(由MCIN/AEI/10.13039/501100011033资助的项目PID2022-137275NA-I00与PID2021-122209OB-C32)、西班牙埃斯特雷马杜拉自治区(项目IDA3-19-0001-3、GR21017与GR21057)以及欧盟(欧洲区域发展基金)的支持。




