#PraCegoVer
收藏资源简介:
#PraCegoVer是一个专为葡萄牙语图像字幕生成设计的大型多模态数据集,由巴西坎皮纳斯大学的计算研究所创建。该数据集基于Instagram上的#PraCegoVer标签下的帖子,旨在通过自然语言描述帮助视觉障碍人士更好地理解图像内容。数据集包含超过520,997条记录,每条记录包括图像和对应的葡萄牙语描述。创建过程中,研究团队开发了自动收集和预处理数据的框架,确保数据的质量和多样性。该数据集的应用领域主要集中在提高互联网内容的可访问性,特别是为视觉障碍用户提供服务,同时也支持图像字幕生成技术的研究和发展。
#PraCegoVer is a large-scale multimodal dataset specifically designed for Portuguese image captioning, created by the Institute of Computing at the University of Campinas in Brazil. This dataset is based on posts under the #PraCegoVer hashtag on Instagram, aiming to help visually impaired people better understand image content through natural language descriptions. It contains over 520,997 records, with each record including an image and its corresponding Portuguese caption. During the dataset creation process, the research team developed a framework for automated data collection and preprocessing to ensure the quality and diversity of the dataset. Its main application scenarios focus on improving the accessibility of internet content, especially providing services for visually impaired users, and it also supports research and development of image captioning technologies.




