The CRyPTIC Consortium Dataset
收藏资源简介:
This dataset processes all the raw genetics (FASTQ) files using a Mycobacterial pipeline as implemented in an online cloud platform. Whilst the bioinformatics components are similar (e.g. Clockwork remains the variant caller), there are some differences. This version includes all samples for which we expect to have WGS and pDST data. This dataset contains the high-level data tables produced by the CRyPTIC Consortium. It contains information on a large number of M. tuberculosis complex samples that were collected and collated by the project. In total 53,897 samples have both WGS and pDST data. An additional 11,945 samples only have pDST data. Due to the size of some of the data tables, the larger ones are stored as PyArrow parquet files. These can be e.g. loaded using pandas but one ordinarily needs to first install pyarrow using pip.



