Dataset related to the article " Plasma-based genetic characterization in advanced non-small cell lung cancer: from real-world evidence to prognostic stratification"
收藏资源简介:
This record contains original data used in the article " Plasma-based genetic characterization in advanced non-small cell lung cancer: from real-world evidence to prognostic stratification" to develop a linear support vector machine (SVM) classifier for the binary classification of copy number alteration (CNA) profiles as stable (SCP) or unstable (UCP). We retrospectively evaluated the results of plasma NGS analysis performed at our Institution on advanced non-small cell lung cancer patients by using the AVENIO ctDNA Expanded Kit (Roche Diagnostics, Basel, Switzerland), a panel of 77 genes, able to detect the major classes of genetic alterations. Binary classification into stable/unstable was initially performed by visual inspection of individual CNA profiles by two independent professionals of our group and used as target. To automatically classify profiles as SCP or UCP beyond operators’ experience, we implemented a support vector machine (SVM) classifier. Briefly, we considered the segmented log2 ratios (.cns) files from the output of AVENIO Oncology Analysis Software and computed three features (Segments, Size, Chromosomes). An alteration in the CNA profile was defined as a DNA segment of any size with | log2 copy ratio | > cut-off. A cut-off of 0.1 was selected and three features were considered as covariates in the SVM classifier: 1) number of altered segments (Segments), 2) total length of altered regions (Size) and 3) number of affected chromosomes (Chromosomes). All anonymized samples processed at our Institute between 2020 and 2024 (n = 428), including those performed as external service, were used to increase reliability and robustness of the model. Available samples were obtained with two different versions of AVENIO ctDNA Expanded kit (Expanded or Expanded V2) and were kept separated in the machine learning prediction as Dataset 1 (n = 298) and Dataset 2 (n = 130). The “Dataset_1.txt” and “Dataset_2.txt” files are the original data matrices corresponding to the two datasets. Rows represent available samples (n=298 for Dataset 1; n=130 for Dataset 2). Columns contain the following variables: anonymized sample IDs (Sample), the class, “stable” or “unstable”, as assigned by two independent professionals of our group (Class), the corresponding binary label (0 for “stable”, 1 for “unstable”), the three features used as covariates in the SVM classifier and computed as described above (Segments, Size, Chromosomes). For the detailed results of our work, please refer to the full article.



