遇见数据集

Data from analysis of fate of horizontally acquired genes

收藏
Zenodo2026-02-13 更新2026-05-26 收录
官方服务:

资源简介:

Descriptions of files generated in this study:- `2_members_subset.tsv` : corresponding to `members.tsv` file from EggNOG database, but only containing the subset of gene families included in our analysis (i.e., those with at least 2000 genes and at least 10 taxa in each gene family). Contains no headers since the original TSV file did not. The tab-separated columns are `taxonomic ID` (NCBI taxonomic ID for bacteria is "2"), `NOG ID` (the identifier for the gene family in EggNOG database), `number of genes`, `number of taxa`, `comma-separated list of gene IDs`, and `comma-separated list of taxonomic IDs`. - `2_taxonomicGroupings.csv` : Taxonomic grouping information for taxa in our dataset, extracted from NCBI taxonomy database. - `2_PPI.csv`: PPI and COG mapping data for genes in the dataset, extracted from STRING database. - `phylumID_log20_sit80.transfers.csv` : the main CSV output file of the analysis, containing all the inferred inter-phylum HGT events with their associated information. Inferred `direction` column of the transfer is "1" if it is from `taxa A` to `taxa B`, "-1" if it is from `taxa B` to `taxa A`, and "0" if the direction could not be determined. In the filename `log20` refers to the logging level for python logging module (20 corresponds to INFO level instead of DEBUG level), and `sit80` refers to the filtering sequence identity threshold (80%) used for downstream analyses of sequence identities across the entire dataset. Other headers are self-explanatory, and the file is comma-separated. - `classID_log20_sit80.transfers.csv` and `orderID_log20_sit80.transfers.csv` are similar files for class and order-level HGT pairs respectively.- `hgt_zscores.csv` : contains the z-scores of the sequence identities, for each inferred inter-phylum HGT pair of genes, with respect to the distribution of sequence identities of all non-HGT pairs of genes between the same two phyla in the same gene family (NOG). - `hgt_zscores_skipped.csv` : information of inter-phylum HGT pairs from about 6% of the total dataset of gene families, for which the z-scores could not be calculated due to insufficient data (e.g., too few non-HGT pairs of genes between the same two phyla in the same gene family). For the sake of computational efficiency, once we don't find enough data to calculate the z-scores for a gene family for a pair of phyla, we skip the calculation of z-scores for all HGT pairs in that gene family for that phylum pair (marked as `invalid_phylum_pair_cached` in the file). - `proteinid_assembly_mapping.csv` : a mapping file for inter-phylum HGT pairs, that links each protein ID (corresponding to the gene IDs in our dataset) to its corresponding assembly ID, which can be used to retrieve additional information about the genome from which the protein was derived. We used this to check the completeness of the genomes involved in the inferred HGT events. - `work_dir.tar.gz` file contains the entire work directory of the analysis, including all the intermediate files, logs, and results of each step. ------------------------------------------------- For further details on how the files were generated, please refer to the `README.md` file in the code repository of this project (link). The `notebooks` directory in the code repository contains the Jupyter notebooks that were used to analyze these files.For details on the structure of the work directory and the files contained in it, please refer to the readme file in the code repository of this project. Scripts used in this study (apart from the notebooks) are the same as those in the `misc` directory of the code repository, with documentation of how those scripts were run also present in the readme file.

提供机构:
Zenodo
创建时间:
2026-02-13
二维码
社区交流群
二维码
科研交流群
商业服务