srpone/hm-eval
收藏资源简介:
H&M时尚评估数据集是一个用于时尚图像-文本检索的评估基准,基于公开的H&M个性化时尚推荐目录构建。该数据集包含105,100个产品(包括图像和结构化文本)和2,000个零样本查询,支持文本到图像、图像到文本和文本到文本检索任务。数据遵循BEIR风格结构,便于与ZooClaw-Fashion等其他基准使用相同代码路径进行评估。查询通过两阶段管道生成:首先从产品名称和随机属性(如颜色、图案、人口统计类别)组合,然后使用Gemma-4-31B-IT模型重写为自然搜索查询。数据集主要用于测试跨分布泛化能力,特别是在单品牌、大语料库零售目录场景下,是zoodata.ai和ZooClaw平台数据代理基准的一部分。
The H&M Fashion Evaluation Dataset is an evaluation benchmark for fashion image-text retrieval, constructed based on the public H&M personalized fashion recommendation catalog. This dataset contains 105,100 products (including images and structured text) and 2,000 zero-shot queries, supporting text-to-image, image-to-text, and text-to-text retrieval tasks. The dataset follows the BEIR-style structure, enabling compatible evaluation with other benchmarks such as ZooClaw-Fashion via identical code pipelines. Queries are generated through a two-stage pipeline: first, combinations of product names and random attributes (e.g., color, pattern, demographic categories) are created, then rewritten into natural search queries using the Gemma-4-31B-IT model. The dataset is primarily used to test cross-distribution generalization capabilities, especially in single-brand, large-corpus retail catalog scenarios, and is part of the data agent benchmarks for the zoodata.ai and ZooClaw platforms.




