Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-Hi
收藏资源简介:
Movie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-HiDescriptionMovie_Scripts_with_Age_Ratings_by_Pratik_Kalamkar-Hi is a curated collection of 203 Hindi-language full-length movie scripts annotated with official Central Board of Film Certification (CBFC) Age_Rating labels. The dataset was developed to support research in Hindi Natural Language Processing (NLP), Automatic Movie Certification, Content Moderation, Regulatory Artificial Intelligence, Explainable AI, and Long-Document Classification. Unlike many publicly available Hindi-language datasets that focus on reviews, news articles, social media posts, or subtitles, this dataset contains complete movie scripts. The scripts preserve dialogues, scene descriptions, character interactions, narrative flow, and contextual information that are essential for understanding age-sensitive content and certification decisions. Data CollectionMovie scripts were collected from multiple publicly available screenplay repositories, including filmcompanion, Scripts.com, and several additional online script archives and movie-script websites. Some scripts are also made with help of subtitles, by removing timestamps. Since no single source provided both script content and reliable Age_Rating metadata, scripts were gathered from multiple sources and consolidated into a unified dataset. The objective was to create a benchmark resource linking Hindi movie scripts with verified CBFC (India) Age_Rating labels suitable for machine learning and NLP research. Age_Rating Annotation and VerificationObtaining reliable Age_Rating information for Hindi movies required a dedicated annotation and verification workflow. Movie titles and release years were extracted, normalized, and cross-validated using multiple movie metadata services, including: OMDb (Open Movie Database) TMDb (The Movie Database) Supplementary web-based verification sources Automated matching procedures were used to verify title-year consistency and Age_Rating information. Ambiguous records, duplicate titles, remakes, and conflicting metadata were manually reviewed and corrected wherever necessary. This semi-automated verification process improved annotation quality and ensured greater confidence in the final Age_Rating assignments. Dataset StructureEach filename contains: Official CBFC Age_Rating Movie Title Release Year Filename format: Age_Rating_Movie_Title_Year.txt Examples: UA_Dangal_2016.txt A_Gangs_of_Wasseypur_2012.txt U_Taare_Zameen_Par_2007.txt This structure allows researchers to directly derive Age_Rating labels from filenames without requiring additional annotation files. Dataset StatisticsTotal Scripts: 203 Age_Rating Distribution: U: 52 UA: 93 A: 58 Language: Hindi Document Type: Full-Length Movie Scripts Certification Framework: CBFC Age_Rating System Potential Research ApplicationsThis dataset can support research in: Hindi NLP Automatic Age_Rating Prediction Movie Certification Systems Long-Document Classification Content Moderation Violence Detection Profanity Detection Explainable AI Regulatory NLP Transformer-based Long-Context Modeling Cross-Lingual Learning Content Severity Analysis SignificanceHindi is one of the most widely spoken languages in the world, yet publicly available screenplay datasets with verified certification labels remain limited. While movie scripts can often be found online, reliable Age_Rating metadata is rarely available in a structured and research-ready format. This dataset addresses that gap by combining script collection, metadata normalization, automated verification, and manual quality assurance into a single curated resource. The resulting corpus provides researchers with a benchmark for studying certification prediction, content moderation, and long-document understanding in Hindi. CitationIf you use this dataset in academic research, please cite: Kalamkar, P. N., Peddi, P., & Sharma, Y. K. (2026). "Lightweight and Explainable Neural Models for Multilingual Movie Script Certification." International Journal of Information Technology and Computer Science (IJITCS), Volume 18, Issue 2, Pages 146–160. DOI: 10.5815/ijitcs.2026.02.09 @article{kalamkar2026moviecertification, author = {Pratik N. Kalamkar and Prasadu Peddi and Yogesh K. Sharma}, title = {Lightweight and Explainable Neural Models for Multilingual Movie Script Certification}, journal = {International Journal of Information Technology and Computer Science}, year = {2026}, volume = {18}, number = {2}, pages = {146--160}, doi = {10.5815/ijitcs.2026.02.09}}



