This dataset contains various repeat catalogs for the Tandem Repeat Genotyping Tool (TRGT): pathogenic_repeats.hg38.bed contains annotations of 56 known pathogenic repeats. polymorphic_repeats.hg38.be
Each number represents an LCB calculated by MAUVE between CA88 and CO92. Changes between steps are underlined. Negative numbers represent an inverted LCB.
The dataset has four manakins's repeat annotation files. We used Tandem Repeat Finder (TRF, v4.09.1) to identify Tandem repeats. We used de novo methods with RepeatModeler (v2.0.2a) and LTR_Finder (v1
S16 Data. Simple repeats found in reference CLTs. Number (n_rep) of mono- and di-nucleotides (type) found in the genes (Gene) of the reference CLTs (CLT) with a word search of every possible mo