遇见数据集

Movie Metadata and Performance Dataset for Machine Learning Applications

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset contains 5,674 unique movie records and 32 attributes prepared for machine learning, data analysis, and entertainment-domain research. It includes information such as movie title, release year, runtime, rating, vote count, Metascore, gross revenue, genre, certification, director, cast, and textual descriptions. The purpose of this dataset is to support the investigation of relationships between movie characteristics and performance-related indicators. The dataset enables exploration of questions such as how factors including genre, runtime, release year, audience voting behaviour, critical scores, and financial performance are related. It can also be used to investigate whether combinations of movie metadata features can support predictive or classification-based machine learning tasks. The data was prepared from real-world movie metadata through systematic data cleaning, transformation, missing-value treatment, text preprocessing, feature engineering, and statistical standardization. Redundant index artifacts were removed and duplicate records were checked. Missing numerical values were handled using genre-wise median imputation, while missing certification values were treated using categorical imputation. Cleaned and structured versions of genre, director, cast, and description fields were created, along with a primary genre feature. Additional standardized features were generated for movie rating, vote count, Metascore, gross revenue, runtime, and release year using Z-score and T-score transformations. These transformations allow numerical variables with different scales to be more easily compared and analysed. The dataset can be interpreted through exploratory statistical analysis, visualization, correlation analysis, regression, classification, clustering, and feature-based machine learning models. Potential applications include movie rating prediction, genre analysis, popularity and audience engagement analysis, financial performance analysis, recommendation system research, and natural language processing of movie descriptions. The final dataset is intended as a structured research resource for analysing patterns and relationships within movie metadata and for evaluating machine learning approaches using real-world entertainment-domain data.

创建时间:
2026-08-23
二维码
社区交流群
二维码
科研交流群
商业服务