We propose a feature vector approach to characterize the variation in large data sets of biological sequences. Each candidate sequence produces a single feature vector constructed with the number and
aPossible substitution of amino acids into the frame-mutation library (%) according to the given codon degeneracy of primers used for mutagenesis of MG2x1 as shown in the Table S1.