Balanced Hollywood Movies Scripts Age Rating Dataset - 250 by Pratik Kalamkar
收藏资源简介:
# Balanced Hollywood Movies Scripts Age Rating Dataset - 250 by Pratik Kalamkar ## Description The Balanced Hollywood Movies Scripts Age Rating Dataset - 250 by Pratik Kalamkar is a benchmark dataset developed for research in automated movie censorship, age-rating prediction, content classification, ordinal learning, and long-document Natural Language Processing (NLP). The dataset contains 250 full-length English movie scripts balanced across five MPAA age-rating categories: | Age Rating | Scripts ||------------|----------|| G | 50 || PG | 50 || PG-13 | 50 || R | 50 || NC-17 | 50 | Total Scripts: 250 The dataset was created to address the class imbalance commonly observed in real-world movie-rating datasets and to provide a reproducible benchmark for machine learning and deep learning research. ## Dataset Structure Each movie is stored as an individual UTF-8 text file containing the full movie script. The filename convention is: [MPAA_Age_Rating]_[Movie_Title]_[Movie_Year].txt Examples: G_Toy_Story_1995.txtPG_Frozen_2013.txtPG-13_Avatar_2009.txtR_The_Matrix_1999.txtNC-17_Showgirls_1995.txt Each filename is prefixed with the verified MPAA age-rating category and suffixed with the movie release year to facilitate filtering, benchmarking, and reproducible experimentation. The movie scripts were collected from multiple publicly available script repositories, screenplay archives, and movie-related websites. Considerable effort was invested in script collection, verification, organization, age-rating annotation, and dataset balancing. ## Research Motivation Movie age-rating prediction is an inherently ordinal classification problem in which categories possess a natural order of increasing content restriction. Most publicly available movie-script datasets exhibit substantial class imbalance, making fair comparison of classification algorithms difficult. This dataset provides a balanced benchmark corpus that enables consistent evaluation of: - Automated censorship systems- Age-rating prediction models- Ordinal classification methods- Long-document NLP models- Machine learning algorithms- Deep learning approaches- Transformer-based architectures ## Applications - Movie Censorship Prediction- Age Rating Classification- Ordinal Learning- Text Classification- Long Document Analysis- Natural Language Processing- Artificial Intelligence Research ## Dataset Construction The dataset was constructed by collecting full-length movie scripts from multiple publicly available sources (IMDB, ScriptORama, Scripts.com) and associating each script with its verified MPAA age rating and release year. A balanced benchmark corpus was then created by selecting: - 50 G-rated scripts- 50 PG-rated scripts- 50 PG-13-rated scripts- 50 R-rated scripts- 50 NC-17-rated scripts resulting in a total of 250 full-length movie scripts suitable for age-rating prediction and censorship classification research.## Citation If you use this dataset in academic research, please cite: Kalamkar, P. N., & Sharma, Y. K. (2025). Hierarchical Ordinal Framework for Automated Movie Censorship Using Full-Length Scripts. 2025 IEEE 6th Global Conference for Advancement in Technology (GCAT). DOI:https://doi.org/10.1109/GCAT66372.2025.11368510 ## Acknowledgment This dataset was created to provide a balanced benchmark for movie age-rating prediction research. If it saves you time, effort, or frustration, please consider citing the associated publication. IEEE Format P. N. Kalamkar and Y. K. Sharma, "Hierarchical Ordinal Framework for Automated Movie Censorship Using Full-Length Scripts," 2025 IEEE 6th Global Conference for Advancement in Technology (GCAT), 2025, doi: 10.1109/GCAT66372.2025.11368510. BibTeX @INPROCEEDINGS{11368510, author={Kalamkar, Pratik N. and Sharma, Yogesh Kumar}, booktitle={2025 IEEE 6th Global Conference for Advancement in Technology (GCAT)}, title={Hierarchical Ordinal Framework for Automated Movie Censorship Using Full-Length Scripts}, year={2025}, doi={10.1109/GCAT66372.2025.11368510}}



