sedthh/tv_dialogue
收藏资源简介:
--- dataset_info: features: - name: TEXT dtype: string - name: METADATA dtype: string - name: SOURCE dtype: string splits: - name: train num_bytes: 211728118 num_examples: 2781 download_size: 125187885 dataset_size: 211728118 license: mit task_categories: - conversational - text2text-generation - text-generation language: - en tags: - OpenAssistant - transcripts - subtitles - television pretty_name: TV and Movie dialogue and transcript corpus size_categories: - 1K<n<10K --- # Dataset Card for "tv_dialogue" This dataset contains transcripts for famous movies and TV shows from multiple sources. An example dialogue would be: ``` [PERSON 1] Hello [PERSON 2] Hello Person 2! How's it going? (they are both talking) [PERSON 1] I like being an example on Huggingface! They are examples on Huggingface. CUT OUT TO ANOTHER SCENCE We are somewhere else [PERSON 1 (v.o)] I wonder where we are? ``` All dialogues were processed to follow this format. Each row is a single episode / movie (**2781** rows total) following the [OpenAssistant](https://open-assistant.io/) format. The METADATA column contains dditional information as a JSON string. ## Dialogue only, with some information on the scene | Show | Number of scripts | Via | Source | |----|----|---|---| | Friends | 236 episodes | https://github.com/emorynlp/character-mining | friends/emorynlp | | The Office | 186 episodes | https://www.kaggle.com/datasets/nasirkhalid24/the-office-us-complete-dialoguetranscript | office/nasirkhalid24 | | Marvel Cinematic Universe | 18 movies | https://www.kaggle.com/datasets/pdunton/marvel-cinematic-universe-dialogue | marvel/pdunton | | Doctor Who | 306 episodes | https://www.kaggle.com/datasets/jeanmidev/doctor-who | drwho/jeanmidev | | Star Trek | 708 episodes | http://www.chakoteya.net/StarTrek/index.html based on https://github.com/GJBroughton/Star_Trek_Scripts/ | statrek/chakoteya | ## Actual transcripts with detailed information on the scenes | Show | Number of scripts | Via | Source | |----|----|---|---| | Top Movies | 919 movies | https://imsdb.com/ | imsdb | | Top Movies | 171 movies | https://www.dailyscript.com/ | dailyscript | | Stargate SG-1 | 18 episodes | https://imsdb.com/ | imsdb | | South Park | 129 episodes | https://imsdb.com/ | imsdb | | Knight Rider | 80 episodes | http://www.knightriderarchives.com/ | knightriderarchives |
数据集信息: 特征: - 名称:TEXT,数据类型:字符串 - 名称:METADATA,数据类型:字符串 - 名称:SOURCE,数据类型:字符串 数据划分: - 划分名称:训练集,字节数:211728118,样本数:2781 下载大小:125187885 字节 数据集总大小:211728118 字节 许可证:MIT 任务类别: - 对话式任务 - 文本到文本生成 - 文本生成 语言:英语 标签: - OpenAssistant - 转录文本(transcripts) - 字幕(subtitles) - 电视剧(television) 友好展示名称:影视对白与脚本语料库 样本规模类别:1000 < 样本数 < 10000 --- # "tv_dialogue" 数据集卡片 本数据集收录了来自多渠道的知名电影与电视剧的脚本与对白文本。 一段示例对白如下: [角色1] 你好 [角色2] 你好,角色2! 最近怎么样? (二人正在对话) [角色1] 我很乐意成为Huggingface平台上的示例。 他们都是Huggingface平台上的示例。 镜头切换至另一场景 我们现在身处别处 [角色1(旁白)] 我想知道我们在哪儿? 所有对白均已按照上述格式进行标准化处理。数据集每一行对应单集电视剧或一部电影,总计2781行,整体遵循[OpenAssistant](https://open-assistant.io/)格式。其中METADATA列以JSON字符串形式存储额外信息。 ## 仅对白版:附带部分场景信息 | 影视名称 | 脚本数量 | 获取渠道 | 数据集标识 | |----|----|---|---| | 《老友记》 | 236集 | https://github.com/emorynlp/character-mining | friends/emorynlp | | 《办公室(美版)》 | 186集 | https://www.kaggle.com/datasets/nasirkhalid24/the-office-us-complete-dialoguetranscript | office/nasirkhalid24 | | 漫威电影宇宙 | 18部电影 | https://www.kaggle.com/datasets/pdunton/marvel-cinematic-universe-dialogue | marvel/pdunton | | 《神秘博士》 | 306集 | https://www.kaggle.com/datasets/jeanmidev/doctor-who | drwho/jeanmidev | | 《星际迷航》 | 708集 | http://www.chakoteya.net/StarTrek/index.html 基于 https://github.com/GJBroughton/Star_Trek_Scripts/ | statrek/chakoteya | ## 完整脚本版:附带详细场景信息 | 影视名称 | 脚本数量 | 获取渠道 | 数据集标识 | |----|----|---|---| | 热门电影 | 919部 | https://imsdb.com/ | imsdb | | 热门电影 | 171部 | https://www.dailyscript.com/ | dailyscript | | 《星际之门:SG-1》 | 18集 | https://imsdb.com/ | imsdb | | 《南方公园》 | 129集 | https://imsdb.com/ | imsdb | | 《霹雳游侠》 | 80集 | http://www.knightriderarchives.com/ | knightriderarchives |
数据集概述
基本信息
- 名称: TV and Movie dialogue and transcript corpus
- 别名: tv_dialogue
- 语言: 英语 (en)
- 任务类别:
- 对话式
- 文本到文本生成
- 文本生成
- 标签:
- OpenAssistant
- 转录
- 字幕
- 电视
- 许可证: MIT
- 大小类别: 1K<n<10K
数据集结构
- 特征:
- TEXT: 字符串类型
- METADATA: 字符串类型
- SOURCE: 字符串类型
- 分割:
- train: 2781个示例,总字节数211728118
- 下载大小: 125187885字节
- 数据集大小: 211728118字节
内容描述
- 内容类型: 电影和电视节目的对话及转录文本
- 示例格式: 每个对话包含人物标识和对话内容,如
[PERSON 1] Hello - 数据量: 总共2781个条目,每个条目代表一个剧集或电影
- 元数据: METADATA列包含JSON格式的额外信息
数据源
- 电视剧集:
- Friends: 236 episodes
- The Office: 186 episodes
- Doctor Who: 306 episodes
- Star Trek: 708 episodes
- 电影:
- Marvel Cinematic Universe: 18 movies
- 其他:
- Top Movies: 919 movies
- South Park: 129 episodes
- Knight Rider: 80 episodes




