FORGE
收藏资源简介:
FORGE是一个为生成式检索构建语义标识符的综合基准,使用来自中国最大的电子商务平台淘宝的用户行为和物品的多模态特征。该数据集包含14亿用户交互和多模态特征的2.5亿物品,用于探索和验证生成式检索中语义标识符的构建优化策略。FORGE旨在解决当前研究在生成式检索中面临的三个主要挑战:缺乏具有多模态特征的大规模公开数据集、对SID生成优化策略的有限调查,以及工业部署中的在线收敛速度慢。FORGE为研究人员提供了一个包含大量用户交互和多模态物品特征的数据集,以及用于评估SID质量的新指标,并引入了一种离线预训练模式,以加速新SID在生产中的收敛。
FORGE is a comprehensive benchmark for building semantic identifiers (SID) for generative retrieval, which leverages multimodal features of user behaviors and items from Taobao, the largest e-commerce platform in China. This dataset encompasses 250 million items with multimodal features and 1.4 billion user interaction records, which is utilized to explore and validate optimization strategies for SID construction in generative retrieval. FORGE aims to address three core challenges in current generative retrieval research: the absence of large-scale public datasets with multimodal features, limited investigations into optimization strategies for SID generation, and slow online convergence during industrial deployment. FORGE provides researchers with a dataset containing abundant user interaction records and multimodal item features, alongside novel metrics for evaluating SID quality, and introduces an offline pre-training paradigm to accelerate the convergence of new SIDs in production environments.
FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets
基本信息
- 标题: FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets
- arXiv ID: 2509.20904
- 学科类别: Computer Science > Information Retrieval (cs.IR)
- 提交日期: 2025年9月25日
- 版本: v1
- 作者: Kairui Fu, Tao Zhang, Shuwen Xiao, Ziyang Wang, Xinming Zhang, Chenchi Zhang, Yuliang Yan, Junjun Zheng, Yu Li, Zhihong Chen, Jian Wu, Xiangheng Kong, Shengyu Zhang, Kun Kuang, Yuning Jiang, Bo Zheng
摘要
语义标识符(SIDs)因其有意义的语义可区分性在生成式检索(GR)中受到越来越多的关注。然而,当前的SIDs研究面临三个主要挑战:
- 缺乏具有多模态特征的大规模公共数据集。
- 对SID生成优化策略的研究有限,这些策略通常依赖于昂贵的GR训练进行评估。
- 在工业部署中在线收敛速度慢。
为解决这些挑战,提出了FORGE,一个用于在工业数据集的生成式检索中形成语义标识符的综合基准。
数据集与基准详情
- 数据集来源: 从中国最大的电子商务平台之一淘宝采样。
- 数据集规模: 包含140亿用户交互和2.5亿项目的多模态特征。
- 基准目标: 探索多种优化以增强SID构建,并通过不同设置和任务的离线实验验证其有效性。
- 在线分析结果: 在日服务超过3亿用户的平台上进行在线分析,显示交易数量增加了0.35%。
方法论与贡献
- 新评估指标: 提出了两个与推荐性能正相关的SID新指标,无需任何GR训练即可进行便捷评估。
- 实际应用优化: 引入了一种离线预训练方案,将在线收敛时间减少了一半。
- 资源可用性: 代码和数据可在 https://doi.org/10.48550/arXiv.2509.20904 获取。




