遇见数据集

Data from: Accurate detection of cholangiocarcinoma in primary sclerosing cholangitis using DNA methylation biomarkers in bile and plasma

收藏
Zenodo2026-06-09 更新2026-06-12 收录
官方服务:

资源简介:

Description of the data and file structure This dataset consists of raw data from droplet digital PCR (ddPCR) analysis of DNA methylation biomarkers for accurate detection of cholanciocarcinoma (CCA). The analyses were performed in tissue, bile and/or plasma from patients diagnosed with PSC alone (without any malignancy), PSC with concomitant CCA (PSC-CCA), sporadic CCA (no concomitant PSC), and non-malignant liver disease other than PSC (disease controls). In tissue and bile, each biomarker was run separately in a duplex ddPCR reaction (QX200, BioRad) with the 4Plex internal control (Pharo et al. Clin Epigenetics. 2018). In plasma, four biomarkers were multiplexed together with the 4Plex (QX600, BioRad). The raw ddPCR data was exported from the QuantaSoft Software 1.7.4.0917 or the QXManager Software version 2.3 (BioRad) in the form of amplitude csv files, and one csv file represents one ddPCR reaction/well. In the csv files, each line indicates the fluorescence amplitude intensity data from one droplet, and the number of lines corresponds to the number of droplets generated for the respective reaction. For tissue and bile run on QX200: The first column ("Ch1 Amplitude") indicates the fluorescence amplitude value of the biomarker assay (Channel 1; FAM) and the second column ("Ch2 Amplitude") indicates the fluorescence amplitude value of the 4Plex internal control (Channel 2; VIC). The third column ("Cluster") indicates which cluster the droplet has originally been assigned to (i.e. double negative, single positive for either the biomarker or the 4Plex control, or double positive). For plasma run on QX600 (multiplexing): The first six columns (Ch1Amplitude - Ch6Amplitude) indicates the droplet fluorescence amplitude value of each of the 6 available channels. The six last columns indicate if the target (i.e. biomarker) is positive, negative or unclassified in each of the droplets in the respective channels. NB. This scoring is set by default by the software, but is not taken into account in downstream analyses. The current data set consits of 3 main folders (tissue, bile and plasma) with raw ddPCR data from both patient and control samples, as well as and negative and positive controls. One folder contains data from samples run on the same 96-well plate. Below follows an explanation for the data in the 3 main folders: "Tissue": Contains 8 subfolders (Amplitudes_"biomarker-name"), each with data from the analyses of 8 different biomarker candidates in the same set of samples and controls (n=24). The name of the biomarker candidate is indicated in the name of the subfolder. "Bile": Contains five subfolders (Plate 1–5). Each Plate folder contains subfolders named Amplitudes_[biomarker name], and Amplitude subfolders include the same samples for that specific marker. The five plate folders contain different samples, and altogether 460 samples were analyzed across all plates. "Plasma": Contains 4 subfolders (Plate1-4) with data from the 4 most promising biomarkers in the same set of samples and controls (n=160). The samples have been analyzed on 4 different plates (Plate 1-4; name of the subfolders) with all 4 biomarkers (and the 4Plex internal control) in a multiplex reaction. For downstream analysis, all csv files were uploaded platewise (i.e. folderwise) to the partition classification algorithm PoDCall - POsitive Droplet CALLer - which automatically dicthomizes positive and negative droplets and returns both non-normalized and normalized concentrations. PoDCall is an R package with an accompagnying shiny graphical user interface (GUI), and is freely availabel on Bioconductor: Bioconductor - PoDCall (https://www.bioconductor.org/packages/release/bioc/html/PoDCall.html) For further information or questions, the corresponding author, Guro E. Lind, can be contacted.

提供机构:
Zenodo
创建时间:
2026-06-09
二维码
社区交流群
二维码
科研交流群
商业服务