遇见数据集

Raw sequence and Non-B-DNA occurrence datasets

收藏
Figshare2023-05-04 更新2026-04-08 收录
官方服务:

资源简介:

<strong>1.Sequence Data Collection: </strong> We have extracted genomic regions [-500 to +500] relative to the translation start sites for 1180 cellular organisms from various public repositories. Archaea and bacteria datasets were retrieved from the NCBI database. Fungal datasets corresponding to Aspergillus, Candida, and Saccharomyces species were extracted from the Aspergillus Genome Database, the Candida Genome Database, and the Saccharomyces Genome databases<strong>.</strong> The UCSC genome browser was used to retrieve fly, mammalian, vertebrate, and worm species. Plant datasets were downloaded from the plant genome database. 1180 unique species belonging to three domains of life have been classified into 28 taxonomic groups loosely based on NCBI taxonomy. <strong>2. Datasets for Non-B DNA motifs:</strong> We computed six putative non-B DNA forming sequences using regular expression models (APR, DR, GQ, IR, MR, and Z) developed by <em>Cer et al</em>. non-B DNA Motif Search Tool (nBMST) in extracted genomic regions. TSV files derived from sequences of each genome were reported in the dataset. 1. Curved DNA motif (A-Phased Repeat, APR) constitutes 3 or more A-tracts (3–5 As) with 10 nucleotides (nt) pacing in the center of each tract. 2.Slipped DNA motifs were defined as 10–50nt direct repeats (DRs) with no intervening nucleotides. 3.G-quadruplexes (GQs) are screened as 4 or more runs of G-tracts (3–5 G’s) separated by 1–7nt spacers. 4.Cruciform DNA is defined as 10–100nt inverted repeats (IRs) with a small spacer size (0–3nt). 5.The triplex DNA motif constitutes10–100nt sized mirror repeats (MRs) with 0–8nt spacers. The sequence composition must be 90% or more purine or pyrimidine nucleotides. 6. The Z-DNA motif is predicted from G-Y runs where G is followed by Y (C or T) for at least 10nt with a strand having alternating Gs. 7. Short tandem repeats (STRs) were also included in the analysis for comparison with other motifs.

1. 序列数据采集:本研究从各类公共资源库中,提取了1180种细胞生物的、相对于翻译起始位点(translation start sites)的[-500至+500]区间基因组区域。其中,古菌与细菌数据集取自NCBI(美国国家生物技术信息中心,National Center for Biotechnology Information)数据库;对应曲霉属、念珠菌属与酿酒酵母属物种的真菌数据集,分别从曲霉基因组数据库、念珠菌基因组数据库及酿酒酵母基因组数据库中提取;果蝇、哺乳类、脊椎动物与线虫类物种的数据集则通过UCSC(加州大学圣克鲁兹分校,University of California, Santa Cruz)基因组浏览器获取;植物数据集下载自植物基因组数据库。本研究共纳入生命三域的1180个独特物种,并依据NCBI分类系统粗略划分为28个分类群。 2. 非B型DNA基序数据集:本研究采用Cer等人开发的正则表达式模型及非B型DNA基序搜索工具(non-B DNA Motif Search Tool,nBMST),在上述提取的基因组区域中,预测得到6种潜在非B型DNA形成序列(分别为APR、DR、GQ、IR、MR与Z型)。数据集内包含了各基因组序列对应的TSV格式文件。具体各类基序定义如下: 1. 弯曲DNA基序(A-Phased Repeat,APR):由3个及以上的A串(3~5个腺嘌呤核苷酸)构成,每个A串中心间距为10个核苷酸(nt)。 2. 滑动DNA基序:定义为无间隔序列的10~50nt正向重复序列(direct repeats,DRs)。 3. G-四链体(G-quadruplexes,GQs):通过筛选得到由4个及以上的G串(3~5个鸟嘌呤核苷酸)构成、且串间间隔1~7个核苷酸的序列。 4. 十字形DNA:定义为间隔序列较短(0~3nt)的10~100nt反向重复序列(inverted repeats,IRs)。 5. 三链DNA基序:由10~100nt的镜像重复序列(mirror repeats,MRs)构成,且串间间隔0~8个核苷酸,其序列组成中嘌呤或嘧啶核苷酸占比需≥90%。 6. Z型DNA基序:通过G-Y串联序列预测得到,即一条链上存在至少10nt的交替G碱基,且G后紧跟嘧啶碱基Y(胞嘧啶C或胸腺嘧啶T)。 7. 为便于与其他基序进行对比分析,本研究同时纳入了短串联重复序列(short tandem repeats,STRs)进行分析。

提供机构:
Yella, Venkata Rajesh
创建时间:
2023-05-04
二维码
社区交流群
二维码
科研交流群
商业服务