遇见数据集

Processed single cell RNA sequencing data from brown trout immune cells of three different origins

收藏
Zenodo2025-11-09 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains processed data and other resources supporting the analyses described in: Ord, J., Martinez, H.S., Solbakken, M.H., Berezenko, A., Oberhaensli, S., Talker, S., Schmidt-Posthaus, H. and Adrian-Kalchhauser, I., 2025. Single-cell RNA sequencing reveals influences of rearing environment on cellular immunity in brown trout (Salmo trutta). bioRxiv, pp.2025-05. 'Ord2025_trout_cellranger_files.tar.gz' This .tar.gz archive contains the raw filtered barcode matrix files for nine samples, as outputted from cellranger-count following alignment of raw 10x Chromium reads to the Salmo trutta genome. The corresponding metadata for each sample are found in Ord2025_trout_sample_metadata.txt. The code used to run cellranger is found here: https://github.com/jamesord/browntrout_scseq/tree/main/Cell_ranger 'Ord2025_trout_Seurat_151025.full.rds' and 'Ord2025_trout_Seurat_151025.downsampled.rds' These files contain the single cell-level raw ('RNA') and integrated expression values of 2000 variably expressed genes from nine samples, plus associated metadata for those samples: sample ID, sampling batch, and group. Here, group refers to one of three origins: 'wild' fish were caught in the wild and are derived from parents presumed to have also been born and raised the wild, 'mix' fish were born and raised in a hatchery facility but parents were born and raised in the wild, while 'farm' fish were born and raised in the hatchery facility to parents that were also born and raised in the hatchery facility. Code for generating these files from the cellranger counts matrix files can be found here: https://github.com/jamesord/browntrout_scseq/blob/main/Seurat/Seurat_SCT_RPCA_clustering_var.ft.filt_161023.R (initial processing) https://github.com/jamesord/browntrout_scseq/blob/main/Seurat/Seurat_subclustering_clustermarkers_180224.R (establishment of final subclusters) 'Ord2025_trout_pseudobulk_counts.zip' This .zip archive contains pseudobulk read counts (i.e. cluster-level sums) for four of the cell clusters: N1, M1, T5 and B3, respectively. Each file contains the total pseudopulk counts of each gene in each of nine samples (see Ord2025_trout_sample_metadata.txt for metadata). Code for generating the pseudobulk counts is found here: https://github.com/jamesord/browntrout_scseq/blob/main/Seurat/pseudobulk/get_pseudobulk_counts_160624.R 'Ord2025_trout_augmented_annotation_material.zip' Contains various files relating to the augmented gene annotation generated using the mikado pipeline. Specifically, the folders and their contents are: blastpContains top BLAST (amino acid) hits of mikado novel gene transcripts, summarised to the gene level (blastphits_genelevel.txt). Where different transcripts of the same gene returned different best hits, alternate best hits are in the third column. The '_edit' version excludes the third column. Also included is the raw output of blastp containing hits of amino acid sequences to SwissProt proteins (mikado.loci.novloc_noveltrans.swissprot.blastp.outfmt6), and the script used to run BLAST. fasta_sequencesTwo sequence files in FASTA format: (1) complete amino acid sequences of novel gene transcripts (mikado.loci.novloc_noveltrans.aa), and (2) complete cDNA sequences (reverse complement of mRNA) of novel gene transcripts (mikado.loci.novloc_noveltrans.cdna.fasta). gtfThe two small gtf files are derived from the Mikado pipeline and formatted such that they can be appended to the ENSEMBL annotation gtf:- gtf annotations of 'novel' genes (mikado.loci.novloc_noveltrans.gtf.gz)- gtf annotations of novel transcripts of existing reference genes; mikado gene IDs were replaced with corresponding ENSEMBL IDs (mikado.loci.refloc_noveltrans.ENS.gtf.gz)Also included is the Salmo trutta ENSEMBL annotation version fSalTru1.1.104 with the above mikado annotations appended (Salmo_trutta.fSalTru1.1.104.filtered.mikado-augmented.gtf.gz). This is the complete augmented annotation file that was used as input to 'cellranger mkref'. Note that the ENSEMBL gtf was however previously filtered using 'cellranger mkgtf' prior to appending the mikado-derived genes. pannzer2DE_genelevel.txt and GO_genelevel.txt are gene descriptions and GO terms derived from PANNZER2 query of mikado novel gene amino acid sequences, with results summarised to the gene level. Also included are the raw outputs from PANNZER2 (in raw_output folder) and two scripts for summarising the outputs to the gene-level (get_genelevel_DE.sh and get_genelevel_GO.sh). PANNZER2 was run using the web interface. For more details, see: http://ekhidna2.biocenter.helsinki.fi/sanspanz/ Please see the github repository for all scripts used in the augmented gene annotation pipeline: https://github.com/jamesord/browntrout_scseq/tree/main/Augmented_annotation

提供机构:
Zenodo
创建时间:
2025-11-09
二维码
社区交流群
二维码
科研交流群
商业服务