遇见数据集

TSV360: A dataset for training and objective evaluation of text-driven 360-degrees video saliency detection methods

收藏
Zenodo2025-09-09 更新2026-05-26 收录
官方服务:

资源简介:

TSV360 is a dataset for text-driven 360-degrees video saliency detection. It contains textual descriptions and the associated ground-truth saliency maps, for 160 videos (up to 60 seconds long) sourced from the VR-EyeTracking and Sports-360 benchmarking datasets. These datasets cover a wide and diverse range of 360-degrees visual content, including indoor and outdoor scenes, sports events, and short films. We constructed the TSV360 dataset, as follows: We utilized EquiRectangular Projection (ERP) frames and their corresponding ground-truth saliency maps from the original datasets. An algorithm processed these inputs to generate multiple 2D video segments, each centered on different events within the same panoramic scene. For each 2D segment, we extracted and assigned event-specific saliency maps derived from the original ground-truth data. Following, these 2D video segments passed through a state-of-the-art video-language model (LlaVA-Next-7B) to generate textual descriptions that capture the depicted events. Finally, we manually curated the generated content to validate and refine it, resulting in 160 videos in total. For each video, there are multiple triplets of ERP frames, saliency maps, and text descriptions, each corresponding to a different event. Note: We release here only the generated saliency maps and text descriptions. The ERP frames are not included and must be obtained separately from the original datasets (see instructions below). How to obtain videos and ERP frames: Download the original videos from the VR-EyeTracking dataset, by following the instructions here: https://github.com/xuyanyu-shh/VR-EyeTracking or here: https://github.com/mtliba/ATSal/tree/master. The subset of these videos that are included in our TSV360 dataset, can be found here: https://github.com/IDT-ITI/TSalV360/blob/main/dataset/vreyetracking.json Download the frames of the videos belonging to the Sports-360 dataset, by following the instructions here: https://github.com/vhchuong/Saliency-prediction-for-360-degree-video/tree/main. The subset of these videos that are included in our TSV360 dataset, can be found here: https://github.com/IDT-ITI/TSalV360/blob/main/dataset/sports360.json Generated saliency maps and text descriptions: In this Zenodo repository we provide the TSV360_gt.rar archive, which contains a subfolder for each of the 160 videos in the dataset, named using the corresponding video_name (e.g. 001, 002, etc.). Inside each video_name/ folder, there are multiple subfolders named using the format video_name_eventIndex (e.g., 001_0/, 001_1/, etc.), where each subfolder corresponds to a distinct event within the 360-degree video, containing saliency maps depicting that event. Each distinct event subfolder contains: description.txt: a textual description of the specific event. .png files: the generated saliency map images corresponding to that event (e.g., 0000.png, 0001.png, etc.), representing sequential indices. For more information about data preparation, you can follow the instructions in section “A. TSV360 dataset” on our GitHub page at: https://github.com/IDT-ITI/TSalV360

提供机构:
Zenodo
创建时间:
2025-09-09
二维码
社区交流群
二维码
科研交流群
商业服务