Structurally Complex with Additive paRent causalitY (SCARY) Dataset
收藏资源简介:
SCARY数据集是由皇家墨尔本理工大学的研究者创建的一个新型合成因果数据集,旨在解决现有数据集在复杂性方面的不足。该数据集包含240个子数据集,每个子数据集有2500个样本,涵盖了40种不同的生成配置,每种配置使用三种不同的种子。数据集通过两种不同的数据生成机制来模拟父母节点与子节点之间的因果关系,包括线性和混合因果机制。SCARY数据集特别关注于模拟真实世界中的选择偏差、不忠实数据和混杂因素,为因果发现算法提供了一个更为真实和挑战性的测试平台。该数据集的应用领域主要集中在因果发现算法的评估和比较,以及探索不同因果发现算法在复杂性和因果关系变化下的表现。
The SCARY dataset is a novel synthetic causal dataset developed by researchers from RMIT University (Royal Melbourne Institute of Technology), designed to address the limitations of existing datasets in terms of complexity. This dataset consists of 240 sub-datasets, each containing 2500 samples, and covers 40 distinct generation configurations, with each configuration utilizing three distinct random seeds. It simulates the causal relationships between parent nodes and child nodes through two different data generation mechanisms: linear and mixed causal mechanisms. Specifically, the SCARY dataset focuses on simulating real-world selection bias, unfaithful data, and confounders, thus providing a more realistic and challenging testbed for causal discovery algorithms. The primary application domains of this dataset center on the evaluation and comparison of causal discovery algorithms, as well as exploring the performance of various causal discovery algorithms under conditions of varying complexity and shifts in causal relationships.




