keratin-fasta
收藏资源简介:
Human Keratin FASTA Sequences 是一个经过策划的人类角蛋白蛋白质FASTA记录数据集,包含从UniProt衍生的序列元数据和计算出的序列组成描述符。该数据集旨在支持角蛋白分子动力学研究,特别是作为伴随研究《Comparative Molecular Dynamics Characterization of Hair Keratin Unfolding Mechanics》的一部分。数据集包含51条记录,覆盖人类(Homo sapiens)的I型和II型角蛋白类别,其中I型25条,II型26条。蛋白质序列长度范围在394至644个氨基酸之间,中位长度为493个氨基酸。数据集中的记录链接到UniProtKB登录号,并注释了生物体、NCBI分类、基因符号、角蛋白类型、序列长度和来源链接。每个序列存储了详细的字段信息,包括记录ID/基因符号、UniProt登录号和URL、FASTA头字段、规范氨基酸序列、校验和(如MD5、SHA-1)、氨基酸组成计数和频率、分组残基组成描述符(如疏水性、电荷性),以及估计的单同位素分子量。此外,数据集包含序列一致性检查的布尔字段,用于标记源序列差异(例如在KRT17和KRT26中发现的序列字符串和长度字段不一致),以便下游用户过滤或审核。数据适用于生物信息学、蛋白质结构分析、分子动力学模拟和角蛋白相关研究。序列记录基于UniProtKB,结构记录链接到AlphaFold Database条目,两者均遵循CC BY 4.0许可。用户在下游工作中应适当引用UniProt、AlphaFold Database和相关研究手稿。
Human Keratin FASTA Sequences is a curated dataset of human keratin protein FASTA records, containing sequence metadata derived from UniProt and computationally calculated sequence composition descriptors. This dataset is intended to support keratin molecular dynamics research, particularly as a companion resource for the study titled "Comparative Molecular Dynamics Characterization of Hair Keratin Unfolding Mechanics". The dataset comprises 51 records covering Type I and Type II keratins from Homo sapiens, with 25 records for Type I keratins and 26 for Type II keratins. The lengths of the protein sequences range from 394 to 644 amino acids, with a median length of 493 amino acids. Each record is linked to a UniProtKB accession number, and is annotated with organism information, NCBI taxonomy, gene symbol, keratin type, sequence length, and source URL. Detailed field information is stored for each sequence, including record ID/gene symbol, UniProt accession number and URL, FASTA header fields, canonical amino acid sequence, checksums (e.g., MD5, SHA-1), amino acid composition counts and frequencies, grouped residue composition descriptors (e.g., hydrophobicity, charge), and estimated monoisotopic molecular weight. Additionally, the dataset includes boolean fields for sequence consistency checks, which are used to flag discrepancies in source sequences (e.g., inconsistencies between sequence strings and length fields observed in KRT17 and KRT26) to enable downstream users to filter or audit the data. This dataset is applicable to bioinformatics, protein structure analysis, molecular dynamics simulations, and keratin-related research. The sequence records are based on UniProtKB, while structural records are linked to AlphaFold Database entries, both of which are licensed under CC BY 4.0. Users should appropriately cite UniProt, the AlphaFold Database, and the associated research manuscript in their downstream work.
数据集概述:Human Keratin FASTA Sequences
数据集名称:Human Keratin FASTA Sequences
许可证:CC BY 4.0
语言:英文
数据集大小:少于1000条记录
标签:生物学、蛋白质、角蛋白、FASTA、UniProt、人类
数据集摘要
- 记录数量:51条
- 生物体:人类(Homo sapiens,NCBI:txid9606)
- 角蛋白类型统计:I型25条,II型26条
- 序列长度范围:394–644个氨基酸
- 序列长度中位数:493个氨基酸
- 存在序列-字符串差异的记录:KRT17、KRT26
- 存在序列长度-字段/字符串差异的记录:KRT17、KRT26
- 生成时间:2026-06-05T15:18:17.879704+00:00
字段说明
每条记录包含以下字段:
- 记录ID/基因符号
- UniProt登录号和URL
- NCBI分类单元和生物体名称
- 角蛋白类型
- FASTA标头字段
- 蛋白质序列及校验和
- 氨基酸组成、分数和估计分子量
- 文件大小、路径、修改时间和SHA-256
数据集生成背景
该序列集作为研究论文《Comparative Molecular Dynamics Characterization of Hair Keratin Unfolding Mechanics》(作者:Wei Lu, Fabien Leonforte, Markus J. Buehler)的一部分进行整理。数据集包含51个人类角蛋白单体蛋白,涵盖I型和II型角蛋白类别。蛋白质记录链接到UniProtKB登录号,并注释了生物体、分类单元、基因符号、角蛋白类别、序列长度和来源链接。
每条序列存储了解析后的FASTA标头、标准氨基酸序列、校验和字段、氨基酸计数和频率、分组残基组成描述符,以及估计的单同位素分子量。序列一致性检查结果以布尔字段保留,便于下游用户筛选或审计存在来源序列差异的记录。
数据来源与归属
序列记录链接到UniProtKB,结构记录(如有)链接到AlphaFold数据库。UniProt数据库的可版权部分根据CC BY 4.0提供,AlphaFold数据库的数据同样根据CC BY 4.0供学术和商业用途。用户应适当引用UniProt、AlphaFold数据库以及相关的角蛋白展开力学研究论文。
验证说明
上传工作流程会检查以下内容:
- 唯一记录标识符
- 预期的I型/II型覆盖范围
- 模拟输出的预期速度组
- 序列长度
- 力向量长度
- 内部PDB残基/SEQRES一致性
逐记录布尔字段保留已知的来源序列差异标记,而非静默覆盖。




