Textual Visual Semantic Dataset
收藏资源简介:
Textual Visual Semantic Dataset是由加泰罗尼亚理工大学TALP研究中心创建的,旨在通过结合视觉上下文信息来提高自然场景中的文本识别能力。该数据集扩展自公开的COCO-text数据集,增加了场景中的物体和地点信息,以帮助研究人员在文本识别系统中加入文本与场景的语义关系。数据集包含图像中的文本候选、场景中的物体、图像位置标签和文本图像描述,利用现有的先进工具提取这些额外信息。该数据集的应用领域包括视觉辅助和自动驾驶等,旨在解决自然图像中自动检测和识别文本的挑战。
Textual Visual Semantic Dataset was developed by the TALP Research Center of the Universitat Politècnica de Catalunya, aiming to improve text recognition capabilities in natural scenes by incorporating visual contextual information. This dataset is extended from the publicly available COCO-text dataset, with added scene objects and location information to help researchers integrate semantic relationships between text and their surrounding contexts into text recognition systems. It contains text candidates in images, scene objects, image location tags, and text-related image descriptions, with these additional pieces of information extracted using existing state-of-the-art tools. Its application fields include visual assistance, autonomous driving and other scenarios, aiming to address the challenges of automatic text detection and recognition in natural images.




