遇见数据集

Speos: An ensemble graph representation framework to predict core genes for complex diseases (Datasets)

收藏
Zenodo2023-01-10 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

the "data.tar.gz" tarball contains the unprocessed or minimally processed data used by Speos. If you intend to use the framework or want to inspect the data, download this part of the dataset. The "final_datasets.tar.gz" tarball contains tsv-formatted, processed data matrices which are directly used as input features for the ensemble models. There are two tsv-files per disease, one labeled "normal", which contains the p input features alongside the gene identifiers and a column which indicates if the gene is labeled as Mendelian or not, and another file labeled "with_n2v_vectors", which also contains the 100-dimensional vectors produced by Node2Vec so the N2V+MLP method can be reproduced with the exact same parameters. All files contain a header row which describes the column and no index column. The "model_parameters.tar.gz" tarball contains the model parameters for all ensemble models used to produce the candidate genes.

"data.tar.gz" 压缩包包含了Speos所使用的未处理或轻度处理的数据。若你打算使用该框架,或希望检视此数据集,请下载该部分数据。 "final_datasets.tar.gz" 压缩包内含以TSV(制表符分隔值,Tab-Separated Values)格式存储的已处理数据矩阵,可直接作为集成模型的输入特征使用。每种疾病对应两份TSV文件:一份标注为"normal",其中包含p个输入特征、基因标识符,以及一列用于标注该基因是否为孟德尔基因(Mendelian)的字段;另一份标注为"with_n2v_vectors",其中额外包含由节点2向量(Node2Vec)生成的100维向量,以便用户可使用完全一致的参数复现N2V+MLP方法。所有文件均包含用于描述列信息的表头行,且未设置索引列。 "model_parameters.tar.gz" 压缩包内含用于生成候选基因的所有集成模型的参数。

提供机构:
Zenodo
创建时间:
2023-01-10
二维码
社区交流群
二维码
科研交流群
商业服务