遇见数据集

A Single-cell Transcriptomic Sequencing Dataset of Early Female and Male Chicken (Gallus gallus) Embryos

收藏
Figshare2025-02-06 更新2026-04-28 收录
官方服务:

资源简介:

Quality Control of Single-Cell DataRaw sequencing data were processed using SCOPE-tools (v1.4.0) to generate a gene expression matrix. After extracting and correcting barcodes and unique molecular identifiers (UMIs), adapter sequences and poly(A) tails were removed. The trimmed reads were aligned to the chicken reference genome (GRCg6a) using the integrated STAR (v2.7.9a) algorithm in CellRanger (v5.0.0). Gene mapping was performed with featureCounts, followed by UMI correction and quantification to produce a complete gene expression matrix. The processed data were then compiled into a matrix file. The expression matrix was further analyzed using the Seurat (v4.3.0.1) package to ensure data quality. Cells were filtered based on gene count thresholds (min.cells > 3 and min.features > 200). Cells with fewer than 1,000 UMIs or a log10GenesPerUMI value exceeding 0.7 were excluded. Additionally, cells with mitochondrial gene content exceeding 25% were removed. These quality control measures ensured the reliability of downstream analyses.Dimensionality Reduction and Clustering of Single-Cell DataTo reduce technical noise and ensure high data quality, the gene expression matrix was normalized and scaled using the NormalizeData and ScaleData functions in the Seurat package. The FindVariableFeatures function was applied to calculate the mean expression and dispersion for each gene, identifying 2,000 highly variable genes. Principal component analysis (PCA) was then performed on the high-dimensional data, retaining the top 20 principal components. Simulated doublet data were generated to match the expected doublet rate, and these were integrated with the original dataset. Each cell was assigned a doublet score using a k-nearest neighbor (k-NN) classifier. Potential doublets were identified using the doubletFinder_v3 function with the parameter pN = 0.25 and removed based on the expected doublet threshold, resulting in a final dataset of 70,361 high-quality cells for downstream analyses. To correct for potential batch effects, the Harmony algorithm was applied. For clustering, the FindClusters function was used with a resolution of 0.4, followed by dimensionality reduction and visualization using uniform manifold approximation and projection (UMAP) and t-distributed stochastic neighbor embedding. The UMAP algorithm was optimized with a neighborhood size of 20 to achieve optimal cell clustering and clear visual representation of the cell populations.Differential Gene Screening To characterize the functional properties of different cell clusters, we identified differentially expressed genes (DEGs) using the "FindAllMarkers" function in the Seurat package. The selection criteria required genes to be expressed in more than 25% of cells in the target cell subpopulation (min.pct = 0.25) and to exhibit significantly higher expression levels in the target cluster compared to others (test.use = "MAST"). To ensure the biological relevance of the results, more stringent thresholds were applied: p-value 1. Cell types were annotated by integrating literature-supported evidence and classical marker genes, allowing for accurate classification of cell populations and elucidation of their biological functions. The expression patterns of marker genes were visualized using the DoHeatmap, DotPlot, and VlnPlot functions in the Seurat package. These visualizations further clarified cell identities and highlighted their functional characteristics.

单细胞数据质量控制 原始测序数据通过SCOPE-tools(v1.4.0)进行处理,以生成基因表达矩阵。提取并校正细胞条形码(barcodes)与唯一分子标识符(UMIs,Unique Molecular Identifiers)后,移除了测序接头(adapter sequences)与poly(A)尾(poly(A) tails)。利用CellRanger(v5.0.0)中整合的STAR算法(v2.7.9a),将修剪后的读段比对至鸡参考基因组(GRCg6a)。通过featureCounts完成基因定位,随后进行UMI校正与定量,以得到完整的基因表达矩阵。随后将处理后的数据整合为矩阵文件。利用Seurat(v4.3.0.1)包对表达矩阵进行进一步分析,以保障数据质量。根据基因计数阈值筛选细胞:min.cells > 3且min.features > 200。排除UMI计数少于1000或log10GenesPerUMI值超过0.7的细胞。此外,移除线粒体基因占比超过25%的细胞。上述质量控制措施保障了后续分析的可靠性。 单细胞数据的降维与聚类 为降低技术噪声并保障数据高质量,我们利用Seurat包中的NormalizeData与ScaleData函数对基因表达矩阵进行归一化与缩放处理。通过FindVariableFeatures函数计算每个基因的平均表达量与离散度,共识别出2000个高可变基因。随后对高维数据进行主成分分析(PCA,Principal Component Analysis),保留前20个主成分。生成与预期双细胞(doublet)比率匹配的模拟双细胞数据,并将其与原始数据集整合。通过k近邻(k-NN)分类器为每个细胞赋予双细胞评分。采用设置参数pN=0.25的doubletFinder_v3函数识别潜在双细胞,并根据预期双细胞阈值移除这些细胞,最终得到70361个高质量细胞用于后续分析。为校正潜在的批次效应,我们采用Harmony算法进行处理。聚类分析采用分辨率为0.4的FindClusters函数完成,随后通过均匀流形近似与投影(UMAP,Uniform Manifold Approximation and Projection)与t分布随机邻域嵌入(t-SNE,t-distributed Stochastic Neighbor Embedding)进行降维与可视化。我们将UMAP算法的邻域大小优化为20,以实现最优的细胞聚类效果,并清晰呈现细胞群体的可视化分布。 差异基因筛选 为阐释不同细胞簇的功能特性,我们利用Seurat包中的FindAllMarkers函数识别差异表达基因(DEGs,Differentially Expressed Genes)。筛选标准要求:基因在目标细胞亚群中超过25%的细胞内表达(min.pct=0.25),且相较于其他细胞簇,在目标簇中呈现显著更高的表达水平(test.use="MAST")。为保障结果的生物学相关性,我们采用更为严格的阈值:p值为1。通过整合文献支持的证据与经典标记基因完成细胞类型注释,实现细胞群体的精准分类并阐明其生物学功能。利用Seurat包中的DoHeatmap、DotPlot与VlnPlot函数对标记基因的表达模式进行可视化,这些可视化结果进一步明确了细胞身份,并凸显了其功能特征。

创建时间:
2025-02-06
二维码
社区交流群
二维码
科研交流群
商业服务