遇见数据集

TSV360: A dataset for training and objective evaluation of text-driven 360-degrees video saliency detection methods

收藏
Zenodo2025-09-09 更新2026-05-26 收录
官方服务:

资源简介:

TSV360 is a dataset for text-driven 360-degrees video saliency detection. It contains textual descriptions and the associated ground-truth saliency maps, for 160 videos (up to 60 seconds long) sourced from the VR-EyeTracking and Sports-360 benchmarking datasets. These datasets cover a wide and diverse range of 360-degrees visual content, including indoor and outdoor scenes, sports events, and short films. We constructed the TSV360 dataset, as follows: We utilized EquiRectangular Projection (ERP) frames and their corresponding ground-truth saliency maps from the original datasets. An algorithm processed these inputs to generate multiple 2D video segments, each centered on different events within the same panoramic scene. For each 2D segment, we extracted and assigned event-specific saliency maps derived from the original ground-truth data. Following, these 2D video segments passed through a state-of-the-art video-language model (LlaVA-Next-7B) to generate textual descriptions that capture the depicted events. Finally, we manually curated the generated content to validate and refine it, resulting in 160 videos in total. For each video, there are multiple triplets of ERP frames, saliency maps, and text descriptions, each corresponding to a different event. Note: We release here only the generated saliency maps and text descriptions. The ERP frames are not included and must be obtained separately from the original datasets (see instructions below). How to obtain videos and ERP frames: Download the original videos from the VR-EyeTracking dataset, by following the instructions here: https://github.com/xuyanyu-shh/VR-EyeTracking or here: https://github.com/mtliba/ATSal/tree/master. The subset of these videos that are included in our TSV360 dataset, can be found here: https://github.com/IDT-ITI/TSalV360/blob/main/dataset/vreyetracking.json Download the frames of the videos belonging to the Sports-360 dataset, by following the instructions here: https://github.com/vhchuong/Saliency-prediction-for-360-degree-video/tree/main. The subset of these videos that are included in our TSV360 dataset, can be found here: https://github.com/IDT-ITI/TSalV360/blob/main/dataset/sports360.json Generated saliency maps and text descriptions: In this Zenodo repository we provide the TSV360_gt.rar archive, which contains a subfolder for each of the 160 videos in the dataset, named using the corresponding video_name (e.g. 001, 002, etc.). Inside each video_name/ folder, there are multiple subfolders named using the format video_name_eventIndex (e.g., 001_0/, 001_1/, etc.), where each subfolder corresponds to a distinct event within the 360-degree video, containing saliency maps depicting that event. Each distinct event subfolder contains: description.txt: a textual description of the specific event. .png files: the generated saliency map images corresponding to that event (e.g., 0000.png, 0001.png, etc.), representing sequential indices. For more information about data preparation, you can follow the instructions in section “A. TSV360 dataset” on our GitHub page at: https://github.com/IDT-ITI/TSalV360

TSV360是一款面向文本驱动360度视频显著性检测的数据集。该数据集包含160段时长不超过60秒的视频的文本描述与对应真实显著性标注图(ground-truth saliency maps),这些视频取材自VR-EyeTracking与Sports-360两个基准数据集。上述基准数据集涵盖了丰富多样的360度视觉内容,包括室内外场景、体育赛事与短片作品。 我们按如下流程构建TSV360数据集:首先从原始数据集中提取等距柱状投影(EquiRectangular Projection, ERP)帧及其对应的真实显著性标注图;随后通过算法对这些输入进行处理,生成多个2D视频片段,每个片段聚焦于同一全景场景内的不同事件;针对每个2D片段,我们从原始真实标注数据中提取并分配针对特定事件的显著性标注图。接下来,我们将这些2D视频片段输入至当前顶尖的视频语言模型LlaVA-Next-7B,以生成能够刻画所描绘事件的文本描述。最后,我们通过人工审核对生成的内容进行验证与优化,最终得到总计160段视频。每段视频均包含多组等距柱状投影帧、显著性标注图与文本描述的三元组,每组对应一个不同的事件。注意:本次仅发布生成的显著性标注图与文本描述,等距柱状投影帧未包含在内,需从原始数据集另行获取(详见后续说明)。 如何获取视频与等距柱状投影帧: 可通过以下链接的说明从VR-EyeTracking数据集中下载原始视频:https://github.com/xuyanyu-shh/VR-EyeTracking 或 https://github.com/mtliba/ATSal/tree/master。TSV360数据集中收录的该数据集子集视频可在此处获取:https://github.com/IDT-ITI/TSalV360/blob/main/dataset/vreyetracking.json 可通过以下链接的说明从Sports-360数据集中下载对应视频的帧:https://github.com/vhchuong/Saliency-prediction-for-360-degree-video/tree/main。TSV360数据集中收录的该数据集子集视频可在此处获取:https://github.com/IDT-ITI/TSalV360/blob/main/dataset/sports360.json 生成的显著性标注图与文本描述:在本Zenodo仓库中,我们提供了TSV360_gt.rar压缩包,其中包含对应数据集160段视频的子文件夹,以对应的视频名称命名(如001、002等)。每个视频名称子文件夹内包含多个以“视频名称_事件索引”格式命名的子文件夹(如001_0/、001_1/等),每个子文件夹对应360度视频中的一个独立事件,内含描述该事件的显著性标注图。每个独立事件子文件夹包含: - description.txt:针对该特定事件的文本描述。 - .png格式文件:对应该事件的生成式显著性标注图(如0000.png、0001.png等),按时序索引命名。 如需了解更多数据集制备相关信息,可参考我们GitHub页面中“A. TSV360数据集”章节的说明:https://github.com/IDT-ITI/TSalV360

提供机构:
Zenodo
创建时间:
2025-09-09
二维码
社区交流群
二维码
科研交流群
商业服务