The CRyPTIC Consortium Dataset
收藏资源简介:
Following the data freeze which produced the original dataset (v1.1.1), additional samples were received from members of the the CRyPTIC Consortium. This version includes these samples. It is not a superset of v1.1.1 since any sample from v1.1.1 whose FASTQ files were downloaded from the ENA are not included -- this reduces the number of genomes significantly, and the number of pDST measurements slightly. Due to this we recommend using later versions if possible as these will be more complete. This dataset contains the high-level data tables produced by the CRyPTIC Consortium. It contains information on a large number of M. tuberculosis complex samples that were collected and collated by the project. In total 44,405 samples with whole-genome sequencing (WGS) information (all derived from short-reads) 56,405 samples with at least one phenotypic drug susceptibility test (pDST) result. Of these 21,570 were collected by CRyPTIC and had the minimum inhibitory concentrations of 13 antibiotics measured using either a UKMYC5 or UKMYC6 96-well broth microdilution plate. 36,738 samples have both WGS and pDST data. Due to the size of some of the data tables, the larger ones are stored as PyArrow parquet files. These can be e.g. loaded using pandas but one ordinarily needs to first pip install pyarrow



