famsa2_tuning_dataset v1.0: reference protein multiple sequence alignments for tuning the dissimilarity measure in FAMSA2
收藏资源简介:
This dataset contains reference protein multiple sequence alignments (MSAs) used to tune the dissimilarity measure in FAMSA2. The data provided here are independent of the benchmarks reported in the manuscript and were used exclusively for parameter selection. Directory structure The dataset contains two main directories: BAliBASE/ - multiple sequence alignment benchmark derived from BAliBASE v3 extHomFam37-2refs/ – extHomFam v37.0 protein families containing exactly two Homstrad reference sequences (the benchmark analyses reported in the manuscript use families with ≥3 reference sequences) Each directory follows the same internal structure: families/ – unaligned protein sequences used as input for alignment [FASTA format] references/ – multiple sequence alignments of reference sequences only, used for accuracy evaluation [FASTA format] Metadata A metadata file (metadata.tsv) provides information on each family: Dataset name - BAliBASE or extHomFam37-2refs Protein family ID - unique identifier for the family (matches the folder name in families/) Total number of sequences - number of sequences in the family Number of reference sequences - number of sequences included in the references/ alignment



