遇见数据集

AMnrGC - Amazon river non-reduntant microbial genes catalogue

收藏
Zenodo2020-09-20 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>AMnrGC : Amazon river basin non-redundant microbial gene catalogue</strong> RELEASE 2018/01<br> -------------------------------------- 1. INTRODUCTION AMnrGC is a collection of genes and proteins which were constructed<br> by use of Amazon river basin openly available metagenomes from<br> sequencing projects (SRP044326, PRJEB25171 and SRP039390). Briefly,<br> metagenomes were coassembled by groups made up their geographical<br> location with Megahit v.1.0 and the contigs were used to gene predictions<br> by Prodigal v.2.6.3. Genes sequences were length filtered (&gt; 150 bp) and<br> clustered by CD-HIT-EST (version 4.6) at 95% of nucleotide identity and<br> 90% of overlap of the shorter gene. Theorical protein products were annotated<br> by the most completes databases up to date and their complete information<br> is available here. 2. LOCATION AMnrGC versions will be available on the web only under the current ZENODO<br> repository: 10.5281/zenodo.1484504 3. FORMAT Gene entries were named as "&gt;AM_AGSSY_XXX" where XXX represents an unique numerical<br> identifier. Genes were deposited in their coding phase, because of this, all of them<br> can be used to generate the protein sequences by transeq function at ORF+1.<br> The protein entries correspond to genes artifical translation used in the annotations,<br> and also available, codified in the same way, but containing the indication "_1" in<br> the end of the header. Example: Gene:<br> &gt;AM_AGSSY_151515 Protein:<br> &gt;AM_AGSSY_151515_1 Annotations were provided as separate tables for each database used to annotate the<br> sequences. The header of these tables indicates the meaning of each value. 2. FUTURE FORMAT CHANGES No major changes are expected for the main general format of the database.<br> New versions should include updated versions of annotations or even additional sequences,<br> numbered as subsequent entries. 3. ACKNOWLEDGEMENTS<br> <br> This work is a joint effort of Laboratory of molecular biology from Federal<br> University of São Carlos, São Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), as well as, the spanish funding organ Consejo Superior de Investigaciones Científicas (CSIC). This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001. 4. THE AMnrGC TEAM AMnrGC is maintained by a group of researchers. You can contact<br> the AMnrGC consortium.<br> <br> Current curators: - Célio Dias Santos Júnior (celio.diasjunior@gmail.com)<br> - Flavio Henrique-Silva (dfhs@ufscar.br)<br> - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> 5. COPYRIGHT NOTICE AMnrGC - Amazon river basin non-redundant microbial gene catalogue<br> Copyright (C) 2018 The AMnrGC consortium. This database is provided “as is” and without any warranty of any kind,<br> of openly available for non-commerical purposes. You can redistribute and/or modify it<br> as you wish, under the terms of the ODBL 1.0 license: https://opendatacommons.org/licenses/odbl/1.0/ For commercial purposes, please contact us. ___________________ The AMnrGC Consortium<br> 2018

**AMnrGC:亚马孙河流域非冗余微生物基因目录(Amazon river basin non-redundant microbial gene catalogue)** 2018年1月发布版 -------------------------------------- 1. 引言 AMnrGC是基于亚马孙河流域公开可用的宏基因组测序项目(SRP044326、PRJEB25171及SRP039390)构建的基因与蛋白质集。简言之,研究团队按地理区域分组,使用Megahit v.1.0对宏基因组进行共组装,并利用Prodigal v.2.6.3对所得重叠群进行基因预测。对基因序列进行长度过滤(保留长度大于150 bp的序列),并使用CD-HIT-EST(版本4.6)按95%的核苷酸同源性、90%的短基因重叠比例进行聚类。基于当前最完备的数据库对理论蛋白产物进行注释,完整注释信息均可在此获取。 2. 数据获取途径 AMnrGC各版本仅可通过当前的ZENODO知识库获取:10.5281/zenodo.1484504 3. 数据格式 基因条目以`>AM_AGSSY_XXX`格式命名,其中XXX为唯一数字标识符。基因以编码区形式存储,因此可通过ORF+1框架下的transeq工具生成对应的蛋白质序列。蛋白质条目为注释所用的基因人工翻译产物,同样以相同格式存储,但序列标题末尾带有`_1`标识。示例如下: 基因条目: `>AM_AGSSY_151515` 蛋白质条目: `>AM_AGSSY_151515_1` 注释信息以单独表格形式提供,对应每一个用于序列注释的数据库,表格表头标注了各列数据的含义。 2. 未来格式变更计划 该数据库的核心通用格式暂无重大变更计划。后续版本将包含更新后的注释版本或新增序列,新增条目将按顺序编号。 3. 致谢 本数据集由巴西圣保罗州圣卡洛斯联邦大学分子生物学实验室(LBM/UFSCAR)与西班牙巴塞罗那海洋科学研究所原生生物研究组(ICM)联合构建。我们感谢巴西国家科学技术发展委员会(CNPq)以及西班牙高等科学研究理事会(CSIC)的资助。 本研究部分受到巴西高等教育人员发展协调局(CAPES)资助,资助编号:001。 4. AMnrGC项目团队 AMnrGC由一批研究人员维护,可联系AMnrGC项目联合体获取相关支持。现任数据管理员: - 塞利奥·迪亚斯·桑托斯·儒尼奥尔(celio.diasjunior@gmail.com) - 弗拉维奥·恩里克-席尔瓦(dfhs@ufscar.br) - 拉米罗·R·洛加雷斯(ramiro.logares@icm.csic.es) 5. 版权声明 AMnrGC——亚马孙河流域非冗余微生物基因目录 版权所有(C)2018 AMnrGC项目联合体。 本数据库按“现状”提供,不附带任何形式的明示或默示担保,仅可用于非商业用途。您可根据ODBL 1.0开源协议(Open Data Commons Open Database License 1.0)的条款重新分发或修改本数据库:https://opendatacommons.org/licenses/odbl/1.0/ 若需用于商业用途,请与我们联系。 ___________________ AMnrGC项目联合体 2018年

提供机构:
Zenodo
创建时间:
2018-11-12
二维码
社区交流群
二维码
科研交流群
商业服务