遇见数据集

CMS-ICH: A Causal-Driven Multimodal Benchmark Dataset for Intangible Cultural Heritage Digital Resource Matching and Propagation Effectiveness Evaluation

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

This dataset, named CMS-ICH (Custom Multimedia Set for Intangible Cultural Heritage), is constructed to support the research of debiased intelligent matching and causal effect optimization for digital cultural resources. It addresses the critical gap in public cultural datasets which lack fine-grained user feedback and multimodal semantic alignment. The data is synthesized based on the statistical distribution constraints defined in the accompanying research paper (Table 2), incorporating a two-level data strategy: (1) using UNESCO and YFCC100M as macro semantic anchors for pretraining, and (2) constructing a high-granularity main benchmark with precise interaction feedback. The dataset comprises 52,480 high-quality cross-modal ICH entities and over 2.31 million simulated behavioral logs collapsed into propagation effectiveness scores. Each sample contains a 512-dimensional dense embedding (embed_0 to embed_511) representing the fused multimodal features (visual, temporal, text, audio, and cultural symbols). It also includes key covariates: user_bias (audience preference), exposure_strength (platform popularity bias), publish_hour, as well as the treatment variable T_treat (intervention intensity) and the ground-truth outcome Y_effect (normalized propagation effectiveness). The true Conditional Average Treatment Effect (CATE_true) is also provided to validate causal debiasing algorithms. This benchmark is specifically designed for evaluating Causal Inference, Double Machine Learning (DML), and Long-tail recommendation algorithms in the cultural heritage domain.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务