MPST
收藏资源简介:
MPST数据集由休斯敦大学计算机科学系创建,包含约14,828部电影的剧情摘要及其关联的71个精细标签。数据集通过MovieLens 20M数据集、IMDb和Wikipedia收集,旨在通过分析电影剧情摘要自动生成标签。数据集的创建过程涉及从多个噪声标签空间中提取相关标签,并通过手动审查和语义相似性聚类来减少标签冗余。该数据集适用于电影标签生成、剧情分析和多标签数据集研究,有助于改进电影推荐系统和观众对电影内容的预先了解。
The MPST dataset was created by the Department of Computer Science at the University of Houston. It contains plot summaries of approximately 14,828 films paired with their associated 71 fine-grained tags. Collected from the MovieLens 20M dataset, IMDb, and Wikipedia, the dataset aims to automatically generate tags by analyzing film plot summaries. The dataset construction process involves extracting relevant tags from multiple noisy label spaces, and reducing tag redundancy through manual review and semantic similarity clustering. This dataset is suitable for research on film tag generation, plot analysis, and multi-label datasets, and it helps improve movie recommendation systems and enhance audiences' prior understanding of film content.



