lamm-mit/keratin-fasta
收藏资源简介:
该数据集名为Human Keratin FASTA Sequences,是一个精心策划的人类角蛋白蛋白FASTA记录集合,包含51条记录,来自人类(Homo sapiens)。数据集涵盖I型角蛋白(25条)和II型角蛋白(26条),序列长度范围为394-644个氨基酸,中位长度为493。每条记录包括UniProt衍生的序列元数据和计算的序列组成描述符,如氨基酸组成、分数和估计分子量。数据集用于支持一项关于头发角蛋白展开力学的比较分子动力学研究,旨在通过模拟分析分子尺度行为如何影响头发纤维的韧性和弹性。数据生成过程涉及从UniProtKB链接蛋白记录,并标注生物体、基因符号、角蛋白类别等信息,同时保留了序列一致性检查标志,以方便下游用户过滤或审核。
This dataset, named Human Keratin FASTA Sequences, is a curated collection of FASTA records for human keratin proteins, containing 51 records derived from Homo sapiens. The dataset covers 25 Type I keratins and 26 Type II keratins, with sequence lengths ranging from 394 to 644 amino acids and a median length of 493. Each record includes UniProt-derived sequence metadata and calculated sequence composition descriptors such as amino acid composition, fractional values, and estimated molecular weight. This dataset supports a comparative molecular dynamics study on the unfolding mechanics of hair keratins, which aims to analyze how molecular-scale behaviors affect the toughness and elasticity of hair fibers through simulations. The data generation process involved retrieving protein records from UniProtKB, annotating them with details including organism, gene symbol, and keratin category, while retaining sequence consistency check flags to facilitate downstream users' filtering and validation.




