BOVText-Benchmark
收藏资源简介:
Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data.In this work, we introduce a large-scale, Bilingual, Open World Video text benchmark dataset (BOVText V2). There are fourfeatures for BOVText V2. Firstly, we provide 2,000+ videos with more than 1,750,000+ frames, 25 times larger than the existinglargest dataset with incidental text in videos. Secondly, our dataset covers 30+ open scenarios, including many virtual scenarios, e.g.,Life Vlog, Driving, Movie, Game, etc. Thirdly, abundant text types annotation (i.e., title, caption or scene text) are provided forthe different representational meanings in the video. Fourthly, the BOVText V2 provides bilingual text annotation to promotemultiple cultures’ lives and communication.
当前绝大多数现有的视频文本检测基准数据集仅聚焦单一语言与场景,且数据规模受限。本工作提出一款大规模双语开放世界视频文本基准数据集(BOVText V2)。该数据集具备四大核心特性:其一,数据集包含2000余段视频,总帧数超175万帧,规模是现有最大视频自然文本数据集的25倍;其二,数据集覆盖30余种开放场景,涵盖诸多虚拟场景,例如生活vlog、驾驶场景、影视片段、游戏画面等;其三,针对视频中承载不同表意功能的文本,数据集提供了丰富的类别标注,涵盖标题、解说字幕与场景文本三类;其四,BOVText V2 采用双语文本标注,以推动多元文化的生活传播与跨文化交流。




