CSE472-blanket-challenge/SCM3K
收藏资源简介:
SCM3K是一个基准数据集,用于评估表格预测中马尔可夫边界的特征选择和预测性能。数据集包含从随机结构因果模型(SCMs)中采样的3,450个表格预测任务,总计345万条记录(每个任务1,000个样本)。每个任务提供了目标节点的真实马尔可夫边界,允许在已知因果结构下评估特征选择和预测。数据集按特征数量分为九个级别(从40到1,000个特征),每个级别对应不同的节点数、DAG密度和马尔可夫边界比例范围。数据生成基于Erdos-Renyi DAGs和六种SCM家族(如线性高斯和非高斯模型),确保了多样性和真实性。数据集适用于机器学习研究,特别是因果推断和特征选择领域。
SCM3K is a benchmark dataset for evaluating feature selection and prediction performance of Markov boundaries in tabular prediction. It consists of 3,450 tabular prediction tasks sampled from random structural causal models (SCMs), totaling 3.45 million records (1,000 samples per task). Each task includes the ground-truth Markov boundary of the target node, enabling evaluation of feature selection and prediction under known causal structure. The dataset is divided into nine feature-count levels ranging from 40 to 1,000, each with varying numbers of nodes, DAG densities, and Markov boundary ratio bands. Data generation uses Erdos-Renyi DAGs and six SCM families (e.g., linear Gaussian and non-Gaussian models), ensuring diversity and realism. The dataset is suitable for machine learning research, particularly in causal inference and feature selection domains.



