遇见数据集

Coding space dimensionality (<i>D</i>) and number (<i>N</i>) of restriction enzymes.

收藏
NIAID Data Ecosystem2026-03-11 收录
官方服务:

资源简介:

The information content in bits, R, of the recognition sequence of 4297 restriction enzymes from REBASE (restriction enzyme database) http://rebase.neb.com or ftp://ftp.neb.com/pub/rebase/ version allenz.801 (Dec 27 2017) [40] was computed. A fully conserved base (A, C, G, T) contributes 2 − log2 1 = 2 bits, two possibilities (R = G/A, Y = C/T, M = A/C, K = G/T, S = C/G, W = A/T) contributes 2 − log2 2 = 1 bit, three possibilities (B = C/G/T, D = A/G/T, H = A/C/T, V = A/C/G) contributes 2 − log2 3 ≈ 0.42 bits and any allowed base (N) contributes 2 − log2 4 = 0 bits [23, 41]. The sum of the information at each base, R, was used to find the corresponding number of compressed bases (λ = R/2) and then the coding dimension (D = 2R), assuming that each enzyme has an efficiency of ϵr = ln 2 and ρ = 1 so that there is a unique dimension according to Eq (21). The most commercially available enzymes and their reported recognition sequences are given as examples. When the DNA backbone cleavage site is known it is indicated by an arrow (↓). The distance to cleavage sites outside the given sequence is shown in parenthesis for the corresponding and complementary strands. Star activity (variation within the canonical site) and flanking sequence effects are found for many restriction enzymes [42]. However, the patterns in the database are reported as consensus sequences that may distort the information content [43], and so may affect the results given here.

创建时间:
2019-10-31
二维码
社区交流群
二维码
科研交流群
商业服务