遇见数据集

Data associated with the reannotation of repeats in the Octopus vulgaris genome assembly (ASM119413v2)

收藏
Zenodo2026-06-16 更新2026-05-26 收录
官方服务:

资源简介:

This is the repeat annotation data generated in "Biological implications of a detailed repeat annotation in Octopus vulgaris" (https://doi.org/10.64898/2026.03.03.709284) for the Octopus vulgaris ASM119413v2 assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001194135.2/). This includes: -the GFF file of the repeat annotation (OctVulg_genome_annotation_only.filteredRepeats.gff) -this is the raw annotation before TE sequences smaller than 100 bp were filtered out for calculating summary information -the FASTA file of reference TE sequences used as well as the new, curated TE and other repeat consensus sequences that were generated (O_vulgaris_and_reference_repeats_Oct24-2.fasta) -sequence headers for reference sequences start with 'REFERENCE', and with 'Ovulg' for new consensus sequences. This includes several consensus sequences of Zinc-finger gene arrays that were noticed during curation (#ZF-array) as well as satellites, unknown repeats and some RNA loci. Characters before the '#' symbol are strings which are unique to each sequence The R markdown file (GBE_O_vulgaris_repeat_annotation.Rmd) contains the code for filtering the GFF file for short TE hits, for young elements and for recreating Figures 1 and 2, including the hotspot/coldspot analysis. Input required for this pipeline is the GFF file of the repeat annotation (OctVulg_genome_annotation_only.filteredRepeats.gff) and the Octopus vulgaris ASM119413v2 assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001194135.2/)

本数据集为发表于论文《Biological implications of a detailed repeat annotation in Octopus vulgaris》(DOI: 10.64898/2026.03.03.709284)的重复序列注释配套数据,针对普通章鱼(Octopus vulgaris)ASM119413v2基因组组装版本(https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001194135.2/)生成。本数据集包含以下内容: - 重复序列注释的GFF文件(OctVulg_genome_annotation_only.filteredRepeats.gff):此为未过滤100bp以下转座因子(Transposable Element, TE)序列的原始注释结果,用于计算汇总统计信息。 - 参考转座因子序列FASTA文件,以及本次研究生成并经人工整理的新型转座因子与其他重复序列共有序列文件(O_vulgaris_and_reference_repeats_Oct24-2.fasta)。 - 参考序列的序列标题以“REFERENCE”开头,新型共有序列的标题则以“Ovulg”开头。本次整理过程中还识别到多个锌指基因阵列(Zinc-finger array, ZF-array)的共有序列,同时包含卫星序列、未知重复序列与部分RNA基因座。#符号前的字符为各序列的唯一标识符。 R Markdown文件(GBE_O_vulgaris_repeat_annotation.Rmd)包含了针对重复序列注释GFF文件的短转座因子匹配结果过滤、年轻元件筛选代码,以及可复现图1与图2的完整代码,其中涵盖热点/冷点分析流程。该分析流程所需的输入文件为重复序列注释GFF文件(OctVulg_genome_annotation_only.filteredRepeats.gff)与普通章鱼ASM119413v2基因组组装文件(https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_001194135.2/)。

提供机构:
Zenodo
创建时间:
2026-05-06
二维码
社区交流群
二维码
科研交流群
商业服务