遇见数据集

Benchmark Data and Deconvolution for Methylome Sequencing: Enriching and Profiling Methylomes for Tumor Classification and Liquid Biopsies

收藏
Zenodo2026-04-08 更新2026-05-26 收录
官方服务:

资源简介:

FLEXseq is a methylation profiling method that targets the adjacent flanks of CCGG motifs in the genome that are associated with cell type-specific markers, enhancers, and other epigenetic functional elements. We benchmarked and demonstrated the versatility of FLEXseq (Fragment Ligation EXclusive methylation sequencing) across plasma cell-free (cf)DNA (P2) and fragmented DNA from formalin-fixed paraffin-embedded (FFPE) tissues (TF725). We sequenced all the samples mentioned above using FLEXseq, whole genome Enzymatic Methyl-seq (EM-seq), and cfDNA reduced representation sequencing with enzymatic conversion (cfDNA-RR-EMseq). Dual-indexed sequencing libraries of EM-seq were constructed with the NEBNext Enzymatic Methyl-seq Kit (E7120L, NEB) with the conversion module (E7125L, NEB), following the manufacturer’s instructions. All libraries above and uniquely dual-indexed libraries with spike-ins of whole genome libraries were pooled and clustered on an Illumina NovaSeqX flow cell with 2×150 bp paired-end sequencing. For the FASTQ files of FLEXseq and EM-seq, after trimming adapters and an extra 100 bp using cutadapt command, paired-end reads were mapped to the human (hg38), lambda, and pUC19 genomes using Bismark. All high-quality sequencing reads were then aligned to the hg38 reference genome using Bismark v0.23.0. We then filtered out reads with unmethylated cytosine in the non-CpG context with filter_non_conversion function. Next, we used the bismark_methylation_extractor function to extract the methylation calls (removing single-nucleotide polymorphisms [SNP]). FLEXseq and EM-seq reads are deduplicated, while cfDNA-RR-EMseq is not due to the same ends of MspI-cut sites. For the deconvolution process, fragment-level deconvolution was used to estimate cell type proportions for each sample with 22 possible cell type references (see Supplementary Methods). Next, we normalized every cell type for each sample as indicated by a z-score. The z-score measures the difference in cell type proportions between target tumors and the reference population. The reference population was defined as either the negative controls or the non-target tumors. The negative control CSF came from patients with autoimmune diseases, infections, or inflammatory conditions (without organ transplants). The z-score was calculated using the statistical equation z-score = (x – μ) / σ. We defined x as the cell type proportion from the case and μ and σ as the mean and standard deviation of the cell type proportion in the reference population, respectively. A z-score of 2 indicates a value of two standard deviations from the mean. Finally, we required the top-ranked cell types to have a z-score > 2 and to be associated with a tumor’s COO (excluding background cell types such as smooth muscle cell and endothelium). The sample was categorized as i) ‘Matched’ when the qualified top-ranked cell type matched the gold standard, ii) ‘Misleading profile’ if it matched a different tumor type, or iii) ‘Indeterminate’ if there was no qualifying top-ranked cell type or if the leading cell type was oligodendrocyte. While oligodendrocytes and neurons had similar neuronal lineages as primary CNS tumors, we used neuron as the COO of CNS tumors. Samples were labeled as ‘Reference’ when used for normalization and ‘N/a’ (not applicable) when excluded for deconvolution classification.

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务