Data and Code for Publication "Testing the Utility of Dental Morphological Trait Combinations for Inferring Human Neutral Genetic Variation"
收藏资源简介:
Data and code for publication: H. Rathmann, H. Reyes-Centeno, Testing the utility of dental morphological trait combinations for inferring human neutral genetic variation. <em>Proc. Natl. Acad. Sci. U.S.A.</em> 117, 10769-10777 (2020). DOI: 10.1073/pnas.1914330117 The repository contains: “R-code.txt”: R code for an exhaustive search algorithm testing the utility of dental morphological traits and trait combinations for inferring human neutral genetic variation. “dental trait frequencies.csv”: Data set with 27 dental morphological trait frequencies for 20 modern human populations worldwide used for analysis. Data from G. R. Scott, C. G. Turner, G. C. Townsend, M. Martinón-Torres, <em>The Anthropology of Modern Human Teeth</em> (Cambridge University Press, 2018). DOI: 10.1017/ 9781316795859 “microsatellite loci mean sizes.csv”: Data set with 645 microsatellite mean allele sizes for 20 modern human populations worldwide used for analysis. Data from T. J. Pemberton, M. DeGiorgio, N. A. Rosenberg, Population structure in a comprehensive genomic data set on human microsatellite variation. <em>G3: Genes Genom. Genet.</em> 3, 891–907 (2013). DOI: 10.1534/g3.113.005728 “utility estimates for 134217727 trait combinations.txt”: A large table with utility estimates for 27 dental morphological traits and all 134,217,700 possible trait combinations. Abbreviations for the 20 population names (rows) in “dental trait frequencies.csv” and “microsatellite loci mean sizes.csv” as follows: AUS = Australia CAS = Central Asia EAF = Eastern Africa EAS = East Asia EEU = Eastern Europe IND = India MAM = Mesoamerica MEL = Melanesia MIC = Micronesia NAF = North Africa NAM = North America NESI = Northeast Siberia NGU = New Guinea NWAM = Na-Dene POL = Polynesia SAM = South America SAN = San SEAS = Southeast Asia WEU = Western Europe WSAF = Sub-Saharan Africa Abbreviations for the 27 dental morphological trait names (columns) in “dental trait frequencies.csv” as follows: T1 = Winging (UI1) T2 = Shoveling (UI1) T3 = Double-Shoveling (UI1) T4 = Interruption Grooves (UI2) T5 = Tuberculum Dentale (UI2) T6 = Mesial Ridge (UC) T7 = Distal Accessory Ridge (UC) T8 = Hypocone (UM2) T9 = Carabelli Trait (UM1) T10 = Cusp 5 (UM1) T11 = Enamel Extensions (UM1) T12 = Peg-Reduced-Missing (UM3) T13 = Lingual Cusp Number (LP2) T14 = Groove Pattern (LM2) T15 = Cusp 6 (LM1) T16 = Cusp Number (LM2) T17 = Deflecting Wrinkle (LM1) T18 = Distal Trigonid Crest (LM1) T19 = Protostylid (LM1) T20 = Cusp 7 (LM1) T21 = Odontomes (UP-LP) T22 = Root Number (UP1) T23 = Root Number (UM2) T24 = Root Number (LC) T25 = Tomes’ Root (LP1) T26 = Root Number (LM1) T27 = Root Number (LM2) Abbreviations for the 645 microsatellite allele locus names (columns) in “microsatellite loci mean sizes.csv” as in T. J. Pemberton, M. DeGiorgio, N. A. Rosenberg, Population structure in a comprehensive genomic data set on human microsatellite variation. <em>G3: Genes Genom. Genet.</em> 3, 891–907 (2013). DOI: 10.1534/g3.113.005728
本数据集与配套代码源自学术论文:H. Rathmann与H. Reyes-Centeno所著《检验牙齿形态特征组合在推断人类中性遗传变异中的效用》(Testing the utility of dental morphological trait combinations for inferring human neutral genetic variation),发表于《美国国家科学院院刊》(Proceedings of the National Academy of Sciences of the United States of America,简称PNAS)117卷,第10769-10777页(2020年),DOI:10.1073/pnas.1914330117。 本代码仓库包含如下内容: 1. "R-code.txt":用于穷举搜索算法的R语言代码,该算法用于检验牙齿形态特征及特征组合在推断人类中性遗传变异中的效用。 2. "dental trait frequencies.csv":包含全球20个现代人群的27项牙齿形态特征频率的数据集,用于本研究分析。数据来源:G. R. Scott、C. G. Turner、G. C. Townsend与M. Martinón-Torres所著《现代人类牙齿人类学》(The Anthropology of Modern Human Teeth),剑桥大学出版社,2018年,DOI:10.1017/9781316795859。 3. "microsatellite loci mean sizes.csv":包含全球20个现代人群的645个微卫星(microsatellite)平均等位基因长度的数据集,用于本研究分析。数据来源:T. J. Pemberton、M. DeGiorgio与N. A. Rosenberg所著《人类微卫星变异综合基因组数据集的群体结构》(Population structure in a comprehensive genomic data set on human microsatellite variation),发表于《G3: Genes, Genomes, Genetics》3卷,第891–907页(2013年),DOI:10.1534/g3.113.005728。 4. "utility estimates for 134217727 trait combinations.txt":包含27项牙齿形态特征及其全部134217700种可能特征组合的效用评估值的大型表格。 ### 人群名称缩写说明 "dental trait frequencies.csv"与"microsatellite loci mean sizes.csv"中20个人群名称(行)的缩写对应如下: AUS = 澳大利亚(Australia)、CAS = 中亚(Central Asia)、EAF = 东非(Eastern Africa)、EAS = 东亚(East Asia)、EEU = 东欧(Eastern Europe)、IND = 印度(India)、MAM = 中美洲(Mesoamerica)、MEL = 美拉尼西亚(Melanesia)、MIC = 密克罗尼西亚(Micronesia)、NAF = 北非(North Africa)、NAM = 北美(North America)、NESI = 东北西伯利亚(Northeast Siberia)、NGU = 新几内亚(New Guinea)、NWAM = 纳-德内语系人群(Na-Dene)、POL = 波利尼西亚(Polynesia)、SAM = 南美(South America)、SAN = 桑人(San)、SEAS = 东南亚(Southeast Asia)、WEU = 西欧(Western Europe)、WSAF = 撒哈拉以南非洲(Sub-Saharan Africa)。 ### 牙齿形态特征缩写说明 "dental trait frequencies.csv"中27项牙齿形态特征名称(列)的缩写对应如下: T1 = 舌侧翼状形态(Winging, UI1)、T2 = 舌侧铲形形态(Shoveling, UI1)、T3 = 双舌侧铲形形态(Double-Shoveling, UI1)、T4 = 舌侧中断沟(Interruption Grooves, UI2)、T5 = 牙结节(Tuberculum Dentale, UI2)、T6 = 近中嵴(Mesial Ridge, UC)、T7 = 远中副嵴(Distal Accessory Ridge, UC)、T8 = 下后尖(Hypocone, UM2)、T9 = 卡氏尖(Carabelli Trait, UM1)、T10 = 第五牙尖(Cusp 5, UM1)、T11 = 釉质延伸(Enamel Extensions, UM1)、T12 = 小牙/缺失(Peg-Reduced-Missing, UM3)、T13 = 舌侧牙尖数(Lingual Cusp Number, LP2)、T14 = 沟型(Groove Pattern, LM2)、T15 = 第六牙尖(Cusp 6, LM1)、T16 = 牙尖总数(Cusp Number, LM2)、T17 = 偏斜皱襞(Deflecting Wrinkle, LM1)、T18 = 远中三角嵴(Distal Trigonid Crest, LM1)、T19 = 原尖脊(Protostylid, LM1)、T20 = 第七牙尖(Cusp 7, LM1)、T21 = 牙瘤(Odontomes, UP-LP)、T22 = 牙根数(UP1)、T23 = 牙根数(UM2)、T24 = 牙根数(LC)、T25 = 托姆斯根(Tomes’ Root, LP1)、T26 = 牙根数(LM1)、T27 = 牙根数(LM2)。 ### 微卫星基因座说明 "microsatellite loci mean sizes.csv"中645个微卫星等位基因座名称(列)的缩写来源详见T. J. Pemberton、M. DeGiorgio与N. A. Rosenberg所著《人类微卫星变异综合基因组数据集的群体结构》,发表于《G3: Genes, Genomes, Genetics》3卷,第891–907页(2013年),DOI:10.1534/g3.113.005728。



