TIGR-100k: a novel benchmark dataset specifically designed for complex query evaluation
收藏资源简介:
TIGR-100k is a novel benchmark dataset designed to advance text-to-image retrieval evaluation by addressing critical limitations in existing benchmarks like COCO and Flickr30k. Unlike traditional datasets that treat retrieval as a "needle-in-a-haystack" task with single ground-truth images, TIGR-100k introduces **compositional complexity** and **multi-relevance grading** to reflect real-world scenarios. Queries are deliberately constructed to include multiple subjects, attributes, or relations (e.g., "a red car parked near a blue house with a dog in the garden"), challenging models to interpret intricate semantics. Crucially, each query is associated with multiple images across three relevance tiers: highly relevant (HR) images that fully match the query, moderately relevant (MR) images differing by only one semantic element (e.g., changing "red car" to "blue car"), and irrelevant (IR) samples. This enables nuanced evaluation beyond binary metrics, assessing a model's ability to retrieve all valid matches, rank semantically proximate alternatives, and handle partial matches. Built on 1,044 bilingual (English and Chinese) complex query pairs, TIGR-100k features 10 HR and 20 MR images per query—synthesized via Stable Diffusion 3.5 Large for precision and semantic variation—yielding approximately 31,320 positive image-query pairs. The full 100k-image corpus integrates these with irrelevant samples from WUKONG, creating a realistic, challenging search space. By emphasizing graded relevance and compositional understanding, TIGR-100k provides a robust framework for evaluating advanced retrieval capabilities, such as fuzzy search and ranking quality, which are essential for practical applications but overlooked in prior benchmarks.



