遇见数据集

Dataset of a Large-Scale English–Arabic Parallel Subtitle Corpus from English-Language Films (2000–2025)

收藏
Mendeley Data2026-07-04 收录
官方服务:

资源简介:

This dataset is a large-scale English–Arabic parallel subtitle corpus compiled from 160 English-language films released between 2000 and 2025. The dataset comprises two complementary files. The first file contains approximately 296,190 aligned English–Arabic subtitle pairs (approximately 2.72 million words), where each record consists of a single English subtitle segment and its corresponding Arabic translation. The second file provides film-level metadata describing each source film used in the corpus. The corpus file includes the following fields: English subtitle segment Arabic subtitle translation Source film title The metadata file contains descriptive information for each film, including: Film title Release year Director(s) Genre British Board of Film Classification (BBFC) rating Runtime (hh:mm) IMDb rating Country of production USA box office revenue Cumulative worldwide box office revenue Number of awards and nominations Production house Source platform (e.g., iTunes, Netflix, Amazon Prime Video, DVD, OSN) The corpus was compiled by collecting publicly available subtitle files, extracting English and Arabic subtitle text, aligning corresponding subtitle segments, and organizing the data into a structured Excel spreadsheet. Basic preprocessing included removing subtitle numbering and timing information, normalizing formatting, verifying subtitle alignment, minimizing duplicate records, and preserving the original subtitle content where appropriate. Each subtitle pair is linked to its source film through the metadata file, enabling linguistic, translational, and film-specific analyses across different genres, production contexts, and release periods. This resource is valuable for: Researchers in audiovisual translation (AVT) and translation studies Corpus linguistics and bilingual corpus research Machine translation and natural language processing (NLP) Terminology extraction and bilingual sentence alignment Arabic–English language technologies Translation and interpreting education Evaluation and benchmarking of language models and subtitle translation systems The dataset supports both qualitative and quantitative investigations of subtitle translation, including translation strategies, lexical and phraseological variation, discourse and pragmatic analysis, cultural adaptation, subtitle segmentation, and corpus-based linguistic research. Furthermore, the accompanying film metadata facilitates analyses examining the influence of genre, production characteristics, release period, and other film-related variables on subtitle translation practices. Together, the two files provide a reusable benchmark resource for developing and evaluating computational and linguistic approaches to English–Arabic subtitle translation.

本数据集为2000年至2025年间上映的160部英语电影构建的大规模英阿平行字幕语料库。本数据集包含两个互补文件:其一包含约296,190对对齐的英阿字幕对(约272万词),每条记录由单条英语字幕片段及其对应阿拉伯语译文构成;其二则提供本语料库所收录每部源电影的电影级元数据。 语料库文件包含以下字段: 英语字幕片段 阿拉伯语字幕译文 源电影标题 元数据文件包含每部电影的描述性信息,包括: 电影标题 上映年份 导演 类型 英国电影分级委员会(British Board of Film Classification, BBFC)评级 片长(格式为hh:mm) IMDb评分 制作国家 美国票房收入 全球累计票房收入 获奖与提名数量 制作公司 来源平台(如iTunes、Netflix、Amazon Prime Video、DVD、OSN) 本语料库通过收集公开可用的字幕文件,提取英阿字幕文本,对齐对应字幕片段,并将数据整理为结构化Excel电子表格构建而成。基础预处理步骤包括移除字幕编号与时序信息、归一化格式、验证字幕对齐、减少重复记录,并在适当情况下保留原始字幕内容。每条字幕对通过元数据文件与其源电影关联,可支持针对不同类型、制作背景与上映周期的语言学、翻译学及电影专项分析。 该资源的适用人群包括: 视听翻译(audiovisual translation, AVT)与翻译研究领域的研究者 语料库语言学与双语语料库研究者 机器翻译与自然语言处理(natural language processing, NLP)从业者 术语提取与双语句子对齐研究人员 英阿语言技术开发者 翻译与口译教育工作者 语言模型与字幕翻译系统的评估与基准测试人员 本数据集支持字幕翻译的定性与定量研究,涵盖翻译策略、词汇与短语变体、语篇与语用分析、文化适配、字幕分段,以及基于语料库的语言学研究。此外,配套的电影元数据可用于分析类型、制作特征、上映周期及其他电影相关变量对字幕翻译实践的影响。综上,两个文件共同构成可复用的基准资源,用于开发与评估英阿字幕翻译的计算与语言学方法。

创建时间:
2026-07-01
二维码
社区交流群
二维码
科研交流群
商业服务