Data from: Two influential primate classifications logically aligned
收藏资源简介:
Classifications and phylogenies of perceived natural entities change in the light of new evidence. Taxonomic changes, translated into Code-compliant names, frequently lead to name:meaning dissociations across succeeding treatments. Classification standards such as the Mammal Species of the World (MSW) may experience significant levels of taxonomic change from one edition to the next, with potential costs to long-term, large-scale information integration. This circumstance challenges the biodiversity and phylogenetic data communities to express taxonomic congruence and incongruence in ways that both humans and machines can process, that is, to logically represent taxonomic alignments across multiple classifications. We demonstrate that such alignments are feasible for two classifications of primates corresponding to the second and third MSW editions. Our approach has three main components: (i) use of taxonomic concept labels, that is name sec. author (where sec. means according to), to assemble each concept hierarchy separately via parent/child relationships; (ii) articulation of select concepts across the two hierarchies with user-provided Region Connection Calculus (RCC-5) relationships; and (iii) the use of an Answer Set Programming toolkit to infer and visualize logically consistent alignments of these input constraints. Our use case entails the Primates sec. Groves (1993; MSW2–317 taxonomic concepts; 233 at the species level) and Primates sec. Groves (2005; MSW3–483 taxonomic concepts; 376 at the species level). Using 402 RCC-5 input articulations, the reasoning process yields a single, consistent alignment and 153,111 Maximally Informative Relations that constitute a comprehensive meaning resolution map for every concept pair in the Primates sec. MSW2/MSW3. The complete alignment, and various partitions thereof, facilitate quantitative analyses of name:meaning dissociation, revealing that nearly one in three taxonomic names are not reliable across treatments—in the sense of the same name identifying congruent taxonomic meanings. The RCC-5 alignment approach is potentially widely applicable in systematics and can achieve scalable, precise resolution of semantically evolving name usages in synthetic, next-generation biodiversity, and phylogeny data platforms.
人类所认知的自然实体的分类与系统发育关系,会随着新证据的出现而发生改变。分类学变动若以符合命名法规的名称形式呈现,往往会在后续的分类处理中引发名称与语义的脱节。诸如《世界哺乳动物物种》(Mammal Species of the World, MSW)这类分类学标准,其各版本间往往会出现显著的分类学变动,这会给长期、大规模的信息整合工作带来潜在的额外成本。这一现状给生物多样性与系统发育数据学界带来了挑战:需要以人类与机器均可处理的方式,表达分类学上的一致性与不一致性,也就是以逻辑化方式表征多种分类体系间的分类学对齐关系。我们以对应《世界哺乳动物物种》第二版与第三版的两类灵长类分类体系为例,证明了此类对齐关系的可行性。我们的方法主要包含三个核心组件:(i) 采用分类学概念标签——即「作者依据的名称」(sec.为secundum的缩写,意为“根据”)——通过父/子层级关系分别构建各概念层级体系;(ii) 通过用户提供的区域连接演算(Region Connection Calculus, RCC-5)关系,明确两个层级体系中选定概念间的关联;(iii) 使用回答集编程(Answer Set Programming)工具包,对这些输入约束进行推理并可视化出逻辑一致的对齐关系。本次研究的用例涵盖两类以Groves的分类体系为依据的灵长类分类:其一为1993年版(对应MSW2,包含317个分类学概念,其中物种级概念233个);其二为2005年版(对应MSW3,包含483个分类学概念,其中物种级概念376个)。通过402条RCC-5输入关联,推理过程最终得到唯一的一致对齐结果,以及153,111条最大信息量关联,这些关联构成了覆盖MSW2与MSW3两套灵长类分类体系中所有概念对的完整语义解析图谱。完整的对齐结果及其各类子集可用于开展名称-语义脱节的量化分析,分析结果显示:近三分之一的分类学名称在不同处理中并不具备可靠性——即同一名称所指代的分类学语义并不一致。RCC-5对齐方法在系统分类学中具备广泛的应用潜力,可在合成型下一代生物多样性与系统发育数据平台中,实现对语义演化的名称用法的可扩展、高精度解析。



