遇见数据集

Data collection for Tsuji et al., 2020, Anoxygenic phototrophic Chloroflexota member uses a Type I reaction center

收藏
Zenodo2023-12-06 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Supplementary data files associated with Tsuji et al., 2020, "Anoxygenic phototrophic Chloroflexota member uses a Type I reaction center". These files are used by code in a corresponding GitHub repository (https://github.com/jmtsuji/Ca-Chlorohelix-allophototropha-RCI) that shows how various analyses that are presented in the paper were conducted. Files included: Capt_S15_sequencer_data_raw.tar.gz -- Gzipped tarball containing the raw Illumina MiSeq output data for the '<em>Candidatus </em>Chlorohelix allophototropha' subculture 15 sequencing run. The run represents a read cloud sequencing run relying on TELL-Seq technology. Indices can be parsed directly from raw output data using the Tell-Read pipeline. scaffold.full.fasta.gz -- the assembled scaffolds generated using Tell-Read and Tell-Link on the above raw MiSeq output data. Ca_Chx_allophototropha_L227-S17_prokka_ORFs.faa.gz and Ca_Chloroheliaceae_bin_L227_5C_prokka_ORFs.faa.gz -- predicted open reading frames (ORFs) from the curated genomes of '<em>Candidatus </em>Chlorohelix allophototropha' and '<em>Candidatus </em>Chloroheliales bin L227-5C', respectively. These ORFs were predicted using prokka and were used for the analyses presented in the paper. They are similar to, but differ somewhat from, the ORFs predicted by the NCBI gene annotation pipeline that was used upon uploading the genomes to the NCBI Genbank database. Thus, these original ORF files are provided for reference in case comparison is ever needed to the publicly accessible Genbank files. I_TASSER_homology_models_full_output.tar.gz -- Gzipped tarball containing the full output from I-TASSER for homology models of key phototrophy-related genes encoded by '<em>Candidatus </em>Chlorohelix allophototropha' and '<em>Candidatus </em>Chloroheliales bin L227-5C'. After unpacking the tarball, view a summary of the I-TASSER output for each gene by clicking on the 'index.html' file in that gene's folder.

本数据集为Tsuji等人2020年发表的题为《不产氧光合绿弯菌门类群利用I型反应中心》的学术论文的配套辅助数据文件。对应GitHub仓库(https://github.com/jmtsuji/Ca-Chlorohelix-allophototropha-RCI)中的代码可复现论文中各项分析的具体实施流程。本次数据集包含以下文件: 1. Capt_S15_sequencer_data_raw.tar.gz:该gzip压缩tar包包含了候选绿弯螺旋菌(Candidatus Chlorohelix allophototropha)15号亚培养物的原始Illumina MiSeq测序产出数据。本次测序采用TELL-Seq技术开展读云测序(read cloud sequencing),其测序索引可通过Tell-Read流程直接从原始测序数据中解析得到。 2. scaffold.full.fasta.gz:该文件为基于上述原始MiSeq测序数据,通过Tell-Read与Tell-Link工具组装得到的基因组支架序列文件。 3. Ca_Chx_allophototropha_L227-S17_prokka_ORFs.faa.gz与Ca_Chloroheliaceae_bin_L227-5C_prokka_ORFs.faa.gz:二者分别为候选绿弯螺旋菌(Candidatus Chlorohelix allophototropha)与候选Chloroheliales类群L227-5C(Candidatus Chloroheliales bin L227-5C)的人工注释基因组所预测的开放阅读框(open reading frame, ORF)。上述开放阅读框通过Prokka工具预测得到,用于支撑论文中的各项分析。其与将基因组上传至NCBI GenBank数据库时采用的NCBI基因注释流程所预测的开放阅读框存在一定差异,但功能相近。为便于后续与公开可获取的GenBank文件进行比对参考,特此提供本批原始开放阅读框文件。 4. I_TASSER_homology_models_full_output.tar.gz:该gzip压缩tar包包含了候选绿弯螺旋菌(Candidatus Chlorohelix allophototropha)与候选Chloroheliales类群L227-5C(Candidatus Chloroheliales bin L227-5C)所编码的关键光合相关基因的同源建模(homology model)完整输出结果,该结果由I-TASSER工具生成。解压该tar包后,可通过点击对应基因文件夹内的`index.html`文件查看该基因的I-TASSER输出摘要。

提供机构:
Zenodo
创建时间:
2020-07-04
二维码
社区交流群
二维码
科研交流群
商业服务