TAAC2025/TencentGR-1M
收藏资源简介:
TencentGR-1M数据集是一个大规模、全模态的数据集,专门为工业广告中的生成式推荐(GR)而设计。该数据集基于真实、去标识化的腾讯广告日志构建,旨在解决GR领域缺乏真实、公开的多模态数据集的问题。数据特征:包含丰富的协同ID和使用最先进嵌入模型提取的多模态表示(文本和视觉)。数据集规模:提供100万条用户序列,每个用户序列最多包含100个交互项目。标签信息:序列中的每个交互都明确标注了曝光(0)和点击(1)信号。
TencentGR-1M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de-identified Tencent Ads logs, it aims to address the lack of realistic, public multi-modal datasets in the GR field. Data Features: Contains rich collaborative IDs and multi-modal representations (text and vision) extracted using state-of-the-art embedding models. Dataset Size: Provides 1 million user sequences, with each user sequence containing up to 100 interacted items. Labels: Each interaction within the sequence is explicitly labeled with exposure(0) and click(1) signals.



