遇见数据集

Data collection for Tsuji et al., 2020, Type I photosynthetic reaction center in an anoxygenic phototrophic member of the Chloroflexota

收藏
Zenodo2023-12-06 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Supplementary data files associated with Tsuji et al., 2021, "Type I photosynthetic reaction center in an anoxygenic phototrophic member of the<em> Chloroflexota</em>". These files are used by code in a corresponding GitHub repository (https://github.com/jmtsuji/Ca-Chlorohelix-allophototropha-RCI) that shows how various analyses that are presented in the paper were conducted. Files included: Capt_S15_sequencer_data_raw.tar.gz -- Gzipped tarball containing the raw Illumina MiSeq output data for the '<em>Candidatus </em>Chlorohelix allophototropha' subculture 15 sequencing run. The run represents a read cloud sequencing run relying on TELL-Seq technology. Indices can be parsed directly from raw output data using the Tell-Read pipeline. scaffold.full.fasta.gz -- the assembled scaffolds generated using Tell-Read and Tell-Link on the above raw MiSeq output data. Ca_Chx_allophototropha_L227-S17_prokka_ORFs.faa.gz and Ca_Chloroheliaceae_bin_L227_5C_prokka_ORFs.faa.gz -- predicted open reading frames (ORFs) from the curated genomes of '<em>Candidatus </em>Chlorohelix allophototropha' and '<em>Candidatus </em>Chloroheliales bin L227-5C', respectively. These ORFs were predicted using prokka and were used for the analyses presented in the paper. They are similar to, but differ somewhat from, the ORFs predicted by the NCBI gene annotation pipeline that was used upon uploading the genomes to the NCBI Genbank database. Thus, these original ORF files are provided for reference in case comparison is ever needed to the publicly accessible Genbank files. I_TASSER_homology_models_full_output.tar.gz -- Gzipped tarball containing the full output from I-TASSER for homology models of key phototrophy-related genes encoded by '<em>Candidatus </em>Chlorohelix allophototropha' and '<em>Candidatus </em>Chloroheliales bin L227-5C'. After unpacking the tarball, view a summary of the I-TASSER output for each gene by clicking on the 'index.html' file in that gene's folder. lake_survey_MAGs.tar.gz -- Gzipped tarball containing the full collection of 756 metagenome-assembled genomes (MAGs) recovered from the Boreal Shield lake survey, corresponding to those mentioned in Supplementary Data 3. The FastA nucleotide genome sequences, FastA nucleotide predicted protein-coding gene sequences, FastA amino acid predicted protein sequences, and Genome Flat Files (GFFs) for all genomes are provided in the fna, ffn, faa, and gff subdirectories, respectively. lake_survey_MAGs_eggnog_annotations.tar.gz -- Gzipped tarball containing annotations (produced via EggNOG) for all predicted proteins among the 756 MAGs recovered from lake metagenome data. Because proteins were pre-clustered prior to annotation, a "orf2gene" file inside the tarball maps the gene clusters to the ORF IDs used for each genome. lake_survey_MAGs_featureCounts.tsv.gz -- GZipped tab-separated table containing the mapping statistics of metatranscriptome reads on all protein-coding genes from the 756 MAGs recovered from lake metagenome data. lake_survey_Ca_Chloroheliales_MAGs_info.tar.gz -- A subset of information from the previous three files specific to genome bins ELA319 and ELA729, which represent RCI-encoding "<em>Ca</em>. Chloroheliales" members.

本数据集为Tsuji等人2021年发表的论文《不产氧光合绿弯菌门(Chloroflexota)成员中的I型光合反应中心》的配套补充数据文件。对应GitHub仓库(https://github.com/jmtsuji/Ca-Chlorohelix-allophototropha-RCI)中的代码可复现论文中各类分析的具体实施流程。 包含的文件如下: 1. Capt_S15_sequencer_data_raw.tar.gz:该gzip压缩打包文件包含“候选绿旋螺菌(Candidatus Chlorohelix allophototropha)”亚培养物15的原始Illumina MiSeq测序产出数据。本次测序采用基于TELL-Seq技术的读云测序方案,测序索引可通过Tell-Read流程直接从原始测序数据中解析。 2. scaffold.full.fasta.gz:基于上述原始MiSeq测序数据,通过Tell-Read与Tell-Link组装得到的基因组支架序列文件。 3. Ca_Chx_allophototropha_L227-S17_prokka_ORFs.faa.gz 与 Ca_Chloroheliaceae_bin_L227_5C_prokka_ORFs.faa.gz:分别为“候选绿旋螺菌”和“候选Chloroheliales菌L227-5C”经人工整理注释的基因组中,通过Prokka预测得到的开放阅读框(open reading frames, ORFs)。这些ORFs用于支撑论文中的相关分析,与将基因组上传至美国国家生物技术信息中心基因银行(NCBI GenBank)数据库时采用的NCBI基因注释流程所预测的ORFs存在相似性,但存在一定差异。为便于后续与公开的GenBank文件进行比对参考,特此提供原始ORF文件。 4. I_TASSER_homology_models_full_output.tar.gz:该gzip压缩打包文件包含针对“候选绿旋螺菌”和“候选Chloroheliales菌L227-5C”所编码的关键光合相关基因的同源建模完整输出结果(由I-TASSER生成)。解压该压缩包后,可通过点击对应基因文件夹内的“index.html”文件查看该基因的I-TASSER输出摘要。 5. lake_survey_MAGs.tar.gz:该gzip压缩打包文件包含从北方盾地湖调查(Boreal Shield lake survey)中获取的756个宏基因组组装基因组(metagenome-assembled genomes, MAGs)的完整集合,对应论文补充数据3中提及的数据集。所有基因组的FASTA格式核苷酸序列、FASTA格式核苷酸编码蛋白基因序列、FASTA格式氨基酸编码蛋白序列以及基因组扁平文件(Genome Flat Files, GFFs)分别存储于fna、ffn、faa与gff子目录中。 6. lake_survey_MAGs_eggnog_annotations.tar.gz:该gzip压缩打包文件包含通过EggNOG对上述756个湖宏基因组数据中所有预测蛋白进行注释得到的结果。由于注释前已对蛋白进行预聚类,压缩包内的“orf2gene”文件可用于将基因簇映射至各基因组对应的ORF标识符。 7. lake_survey_MAGs_featureCounts.tsv.gz:该gzip压缩的制表符分隔文本文件,包含宏转录组测序读段与从湖宏基因组数据中获取的756个MAGs的所有蛋白编码基因的比对统计信息。 8. lake_survey_Ca_Chloroheliales_MAGs_info.tar.gz:该gzip压缩打包文件包含前述三个文件中针对两个基因组分箱(genome bin)ELA319与ELA729的特定子集信息,这两个分箱分别代表编码I型光合反应中心的“候选Chloroheliales”类群成员。

提供机构:
Zenodo
创建时间:
2021-07-24
二维码
社区交流群
二维码
科研交流群
商业服务