EugeneMeng/Hongguo-Short-Drama-Corpus-AI-Labeled
收藏资源简介:
本数据集包含了约1500条来自红果平台的微短剧精选数据,旨在为中文短剧的NLP研究、市场趋势分析以及自动化剧名生成等任务提供高质量的基准数据。数据集具有多维度标注,涵盖了标题、受众、标签、简介及集数,并且使用了AI增强受众标签技术,针对部分原始数据未标注受众标签的问题,使用了专门的sex_divide.py脚本进行预测,并保留了预测置信度。数据字段包括drama_id、title、audience_type、tags、episode_count、description、ai_confidence和label_source。数据来源为红果平台公开信息,标注过程包括提取平台原有的受众标签和使用AI预测缺失标签。
This dataset contains approximately 1,500 selected micro-drama data from the Hongguo platform, aiming to provide high-quality benchmark data for NLP research, market trend analysis, and automated drama title generation tasks for Chinese short dramas. The dataset features multi-dimensional annotations, covering titles, audiences, tags, descriptions, and episode counts, and uses AI-enhanced audience labeling technology. For some original data without audience labels, a specialized sex_divide.py script was used for prediction, and the prediction confidence was retained. Data fields include drama_id, title, audience_type, tags, episode_count, description, ai_confidence, and label_source. The data source is public information from the Hongguo platform, and the annotation process includes extracting original audience labels from the platform and using AI to predict missing labels.





