CMS-ICH: A Causal-Driven Multimodal Benchmark Dataset for Intangible Cultural Heritage Digital Resource Matching and Propagation Effectiveness Evaluation
收藏资源简介:
This dataset, named CMS-ICH (Custom Multimedia Set for Intangible Cultural Heritage), is constructed to support the research of debiased intelligent matching and causal effect optimization for digital cultural resources. It addresses the critical gap in public cultural datasets which lack fine-grained user feedback and multimodal semantic alignment. The data is synthesized based on the statistical distribution constraints defined in the accompanying research paper (Table 2), incorporating a two-level data strategy: (1) using UNESCO and YFCC100M as macro semantic anchors for pretraining, and (2) constructing a high-granularity main benchmark with precise interaction feedback. The dataset comprises 52,480 high-quality cross-modal ICH entities and over 2.31 million simulated behavioral logs collapsed into propagation effectiveness scores. Each sample contains a 512-dimensional dense embedding (embed_0 to embed_511) representing the fused multimodal features (visual, temporal, text, audio, and cultural symbols). It also includes key covariates: user_bias (audience preference), exposure_strength (platform popularity bias), publish_hour, as well as the treatment variable T_treat (intervention intensity) and the ground-truth outcome Y_effect (normalized propagation effectiveness). The true Conditional Average Treatment Effect (CATE_true) is also provided to validate causal debiasing algorithms. This benchmark is specifically designed for evaluating Causal Inference, Double Machine Learning (DML), and Long-tail recommendation algorithms in the cultural heritage domain.



