Benchmarking imputation methods for categorical biological data
收藏资源简介:
Description: Welcome to the Zenodo repository for Publication Benchmarking imputation methods for categorical biological data, a comprehensive collection of datasets and scripts utilized in our research endeavors. This repository serves as a vital resource for researchers interested in exploring the empirical and simulated analyses conducted in our study. Contents: empirical_analysis: Trait Dataset of Elasmobranchs: A collection of trait data for elasmobranch species obtained from FishBase , stored as RDS file. Phylogenetic Tree: A phylogenetic tree stored as a TRE file. Imputations Replicates (Imputation): Replicated imputations of missing data in the trait dataset, stored as RData files. Error Calculation (Results): Error calculation results derived from imputed datasets, stored as RData files. Scripts: Collection of R scripts used for the implementation of empirical analysis. simulation_analysis: Input Files: Input files utilized for simulation analyses as CSV files Data Distribution PDFs: PDF files displaying the distribution of simulated data and the missingness. Output Files: Simulated trait datasets, trait datasets with missing data, and trait imputed datasets with imputation errors calculated as RData files. Scripts: Collection of R scripts used for the simulation analysis. TDIP_package: Scripts of the TDIP Package: All scripts related to the Trait Data Imputation with Phylogeny (TDIP) R package used in the analyses. Purpose: This repository aims to provide transparency and reproducibility to our research findings by making the datasets and scripts publicly accessible. Researchers interested in understanding our methodologies, replicating our analyses, or building upon our work can utilize this repository as a valuable reference. Citation: When using the datasets or scripts from this repository, we kindly request citing Publication Benchmarking imputation methods for categorical biological data and acknowledging the use of this Zenodo repository. Thank you for your interest in our research, and we hope this repository serves as a valuable resource in your scholarly pursuits.
描述: 欢迎使用本存放于Zenodo仓库的数据集,其对应研究论文为《分类生物数据插补方法基准测试》(Benchmarking imputation methods for categorical biological data),本仓库收录了本研究中使用的全套数据集与脚本代码。本仓库可为有意探索本研究实证与模拟分析流程的科研人员提供关键支撑资源。 内容: 实证分析模块: 板鳃类性状数据集:从FishBase获取的板鳃类物种性状数据集合,存储为RDS格式文件。 系统发育树:存储为TRE格式文件的系统发育树。 插补重复样本(插补集):针对性状数据集中缺失值开展的重复插补结果,存储为RData格式文件。 误差计算(结果集):由插补数据集导出的误差计算结果,存储为RData格式文件。 脚本:用于开展实证分析的R脚本集合。 模拟分析模块: 输入文件:用于模拟分析的CSV格式输入文件。 数据分布PDF文档:展示模拟数据与缺失值分布情况的PDF文件。 输出文件:包含模拟性状数据集、含缺失值的性状数据集以及计算了插补误差的插补后性状数据集,所有文件均存储为RData格式。 脚本:用于开展模拟分析的R脚本集合。 TDIP包模块: 本研究分析中使用的系统发育性状数据插补(Trait Data Imputation with Phylogeny, TDIP)R包相关脚本集合。 项目宗旨: 本仓库通过公开数据集与脚本代码,保障本研究成果的透明度与可复现性。有意了解本研究方法、复现本研究分析流程或基于本研究开展后续拓展工作的科研人员,可将本仓库作为重要参考资料。 引用说明: 使用本仓库内的数据集或脚本时,请务必引用本研究论文《分类生物数据插补方法基准测试》(Benchmarking imputation methods for categorical biological data),并注明使用了本Zenodo仓库。 感谢您对本研究的关注,衷心希望本仓库能为您的学术研究提供助力。



