遇见数据集

Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-EN

收藏
Zenodo2026-06-19 更新2026-06-21 收录
官方服务:

资源简介:

Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-ENDescriptionMovie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-EN is a curated collection of 1,142 English-language full-length movie scripts annotated with official Motion Picture Association (MPAA) Age_Rating labels. The dataset was created to support research in Automatic Movie Certification, Content Moderation, Regulatory Natural Language Processing (NLP), Explainable Artificial Intelligence (XAI), and Long-Document Classification. Dataset contains complete movie scripts, preserving dialogues, scene descriptions, narrative structure, contextual information, and thematic progression. Such information closely resembles the material examined by movie certification authorities during age-rating assessment. Data CollectionMovie scripts were collected from multiple publicly available screenplay and script repositories, including IMDb-associated screenplay collections, Scripts.com, and several additional online screenplay archives and script-hosting websites. Since no single source provided both script content and reliable Age_Rating metadata, scripts were gathered from multiple sources and subsequently curated into a unified benchmark dataset. The objective was to construct a large-scale collection of English movie scripts linked with verified age-certification labels suitable for machine learning and NLP research. Age_Rating Annotation and Verification A semi-automated annotation pipeline was developed to associate each movie script with its official MPAA Age_Rating category. Movie titles and release years were extracted, normalized, and cross-validated using multiple movie metadata services, including: OMDb (Open Movie Database) TMDb (The Movie Database) Supplementary web-based verification sources Automated matching procedures were used to verify consistency between script titles, release years, and Age_Rating information. Ambiguous records, duplicate matches, remake conflicts, and metadata inconsistencies were manually reviewed and corrected wherever necessary. This verification process significantly improved the reliability of the final annotations and reduced the likelihood of incorrect script-to-movie mappings. Dataset StructureEach filename contains three key metadata attributes: Official MPAA Age_Rating Movie Title Release Year Filename format: Age_Rating_Movie_Title_Year.txt Examples: R_The_Matrix_1999.txt PG13_Avatar_2009.txt PG_Finding_Nemo_2003.txt This structure enables researchers to directly derive target labels without requiring separate annotation files. Dataset StatisticsTotal Scripts: 1,142 Age_Rating Distribution: G: 65 PG: 153 PG-13: 290 R: 595 NC-17: 39 Language: English Document Type: Full-Length Movie Scripts Certification Framework: MPAA Age_Rating System Potential Research ApplicationsThis dataset can support research in: Automatic Age_Rating Prediction Movie Certification Systems Long-Document Text Classification Content Moderation Violence Detection Profanity Detection Explainable AI Transformer-based Long-Context Modeling Regulatory NLP Content Severity Analysis Narrative Content Analytics SignificanceObtaining movie scripts from public repositories is relatively straightforward; however, constructing a large-scale benchmark with verified Age_Rating annotations with it is considerably more difficult. Certification information is often absent, inconsistent, or distributed across multiple metadata sources. This dataset addresses that challenge by combining script collection, metadata normalization, automated validation, and manual quality assurance into a single curated resource. The resulting corpus provides researchers with a ready-to-use benchmark for studying age-based content classification and regulatory AI systems using full-length movie scripts. CitationIf you use this dataset in academic research, please cite: Kalamkar, P. N., Peddi, P., & Sharma, Y. K. (2026). "Lightweight and Explainable Neural Models for Multilingual Movie Script Certification." International Journal of Information Technology and Computer Science (IJITCS), Volume 18, Issue 2, Pages 146–160. DOI: 10.5815/ijitcs.2026.02.09@article{kalamkar2026moviecertification, author = {Pratik N. Kalamkar and Prasadu Peddi and Yogesh K. Sharma}, title = {Lightweight and Explainable Neural Models for Multilingual Movie Script Certification}, journal = {International Journal of Information Technology and Computer Science}, year = {2026}, volume = {18}, number = {2}, pages = {146--160}, doi = {10.5815/ijitcs.2026.02.09}}

提供机构:
Zenodo
创建时间:
2026-06-19
二维码
社区交流群
二维码
科研交流群
商业服务