Paulamata/Peliculas1940-2025_WikipediaDataset
收藏资源简介:
这是一个由马德里理工大学学生为学术项目创建的数据集,包含1940年至2025年间超过3000部西班牙电影的信息,数据来源于西班牙语维基百科。该数据集旨在提供一个可发现、可访问、可互操作、可重用的FAIR语料库,用于自然语言分析、文本挖掘和电影行业的时间序列分析。数据集结合了文本(剧情简介和标题)、时间变量(发布日期)和结构化元数据(年份、类型、分钟数的时长、剧情简介和唯一ID)。
This dataset has been created by students from Universidad Politécnica de Madrid for an academic project, containing information on over 3000 Spanish movies released between 1940 and 2025, sourced from the Spanish Wikipedia. The dataset aims to provide a FAIR (Findable, Accessible, Interoperable, Reusable) corpus for natural language analysis, text mining, and time-series analysis of the film industry. It combines text (synopsis and title), temporal variables (release date), and structured metadata (year, genre, duration in minutes, synopsis, and unique ID).



