Methylome composite cfDNA classifier (COMET) enhances tumor detection and typing in minimally invasive biopsies
收藏资源简介:
We collected 121 pleural, peritoneal/ascitic, and FNA supernatant fluid specimens consecutively between 2020 and 2021 from Stanford Hospital and followed up until May 2025. We further included 348 body fluid samples from Stanford Hospital between 2021 and 2024, and 62 from the UCSF Clinical Laboratories (San Francisco, CA, USA) between 2020 and 2023. Negative controls included patients who had no cancer history in the past five years and remained cancer-free for at least six months of follow-up. Sequenced data of 106 CSF specimens from previous work[1] were also included in the validation set. A total of 575 specimens were included, and 519 were eventually enrolled according to the inclusion criteria. Those patients with non-CSF specimens were followed up until May 2025. After data filtering, 391 specimens were included. In the controlled FNA experiments, a board-certified cytopathologist (A.C.L) reviewed all cytology smears and provided the diagnostic interpretations. FLEXseq for cfDNA methylome profiling was performed as previously described[1]. Sequencing was performed on an Illumina NovaSeq 6000 or NovaSeqX with either approximately 2x50 bp or 2x150 bp paired-end configurations. Paired-end FASTQ files underwent initial quality control and length trimmed using cutadapt v.4.4. Adapter sequences (NNAGATCGGAAGAGC on both ends) were trimmed, and reads shorter than 25 base pairs were filtered out. The cleaned reads were then aligned to the human reference genome (GRCh38/hg38), lambda, and pUC19 genomes using Bismark v.0.23.0. The final methylation data were deidentified by removing all SNP positions (dbSNP153Common) using bedtools. We used the R-based ichorCNA tool and/or Linux-based CNVkit to estimate the tumor fraction in cfDNA and pellet gDNA. We also used TCGA 450k microarray data as references for t-distributed stochastic neighbor embedding (t-SNE) dimensionality reduction to visualize methylation clustering per sample.



