Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-Mr
收藏资源简介:
Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-MrDescriptionMovie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-Mr is a curated collection of 100 Marathi-language full-length movie scripts annotated with official Central Board of Film Certification (CBFC) Age_Rating labels. The dataset was developed to support research in Marathi Natural Language Processing (NLP), Low-Resource Language Technologies, Automatic Movie Certification, Content Moderation, Explainable AI, and Long-Document Classification. Unlike many publicly available Marathi-language datasets that focus on news articles, social media posts, or short text collections, this dataset contains complete movie scripts. The scripts preserve dialogues, scene descriptions, cultural context, narrative structure, and thematic content that are important for age-certification analysis and content understanding. Data CollectionMovie scripts were collected from multiple publicly available screenplay repositories, including filmcompanion, Scripts.com, and several additional online script archives and movie-script websites. Some scripts are also made with help of subtitles, by removing timestamps. Since no single source provided both script content and verified Age_Rating information, scripts were gathered from multiple sources and consolidated into a unified benchmark dataset. The objective was to create a research resource linking Marathi movie scripts with verified CBFC (India) Age_Rating labels. Age_Rating Annotation and VerificationReliable Age_Rating information was obtained through a semi-automated annotation and validation process. Movie titles and release years were extracted, normalized, and cross-validated using multiple movie metadata services, including: OMDb (Open Movie Database) TMDb (The Movie Database) Supplementary web-based verification sources Automated validation procedures were used to verify title-year consistency and certification information. Ambiguous cases, duplicate movie titles, conflicting metadata, and uncertain matches were manually reviewed and corrected. This process improved annotation reliability and reduced the risk of incorrect script-to-movie associations. Dataset StructureEach filename contains: Official CBFC Age_Rating Movie Title Release Year Filename format: Age_Rating_Movie_Title_Year.txt Examples: U_Sairat_2016.txt UA_Natsamrat_2016.txt UA_Timepass_2014.txt This structure allows researchers to directly derive Age_Rating labels from filenames without requiring separate annotation resources. Dataset StatisticsTotal Scripts: 100 Age_Rating Distribution: U: 50 UA: 50 Language: Marathi Document Type: Full-Length Movie Scripts Certification Framework: CBFC Age_Rating System Potential Research ApplicationsThis dataset can support research in: Marathi NLP Low-Resource Language Modeling Automatic Age_Rating Prediction Long-Document Classification Content Moderation Explainable AI Regulatory NLP Cross-Lingual Learning Transfer Learning Content Severity Analysis Narrative Content Analytics SignificanceMarathi remains underrepresented in Natural Language Processing research compared to major global languages. Publicly available screenplay datasets are particularly scarce, and resources linking full-length scripts with verified Age_Rating information are even rarer. This dataset addresses that gap by providing a curated collection of Marathi movie scripts linked with verified CBFC Age_Rating labels. It offers a benchmark resource for researchers working on low-resource language technologies, content classification, and multilingual regulatory AI systems. CitationIf you use this dataset in academic research, please cite: Kalamkar, P. N., Peddi, P., & Sharma, Y. K. (2026). "Lightweight and Explainable Neural Models for Multilingual Movie Script Certification." International Journal of Information Technology and Computer Science (IJITCS), Volume 18, Issue 2, Pages 146–160. DOI: 10.5815/ijitcs.2026.02.09 @article{kalamkar2026moviecertification, author = {Pratik N. Kalamkar and Prasadu Peddi and Yogesh K. Sharma}, title = {Lightweight and Explainable Neural Models for Multilingual Movie Script Certification}, journal = {International Journal of Information Technology and Computer Science}, year = {2026}, volume = {18}, number = {2}, pages = {146--160}, doi = {10.5815/ijitcs.2026.02.09}}



