Supporting Dataset for the Article "Beyond Interpolation: Integration of Data and AI-Extracted Knowledge for High-Entropy Alloy Discovery"
收藏资源简介:
Discovering novel high-entropy alloys (HEAs) with desirable properties is challenged by the vast compositional space and the complexity of phase formation mechanisms. Several inductive screening methods that excel at interpolation have been developed; however, they struggle with extrapolating to novel alloy systems. This study introduces a framework that addresses the extrapolation limitation by systematically integrating knowledge extracted from material datasets with expert knowledge derived from scientific literature using large language models (LLMs). Central to our framework is the elemental substitution principle, which identifies chemically similar elements that can be interchanged while preserving desired properties. To model and combine evidence from these multi-source knowledge, we employ the Dempster--Shafer theory, which provides a mathematical foundation for reasoning under uncertainty. Our framework consistently outperforms conventional phase selection models that rely on single-source knowledge across all experiments, showing notable advantages in predicting phase stability for compositions containing elements absent from training data. Importantly, the framework intends to effectively complement the strengths of the existing methods. Moreover, it provides interpretable reasoning that elucidates element substitutability patterns critical to alloy stability in HEAs formation. These results highlight the framework's potential for knowledge transfer and extrapolation, offering an efficient approach to exploring the vast compositional space of HEAs with enhanced generalizability and interpretability. Experiments are conducted considering four computational datasets of quaternary alloys, one experimental dataset of quaternary alloys, and one experimental dataset of quinary high-entropy borides (HEB). HEBs are single-phase ceramics containing multiple transition metal cations randomly distributed on the metal sublattice of a boride structure, offering unique combinations of metallic and ceramic properties. Despite different bonding mechanisms, HEBs exhibit similarly high elemental selectivity as HEAs--boron's restrictive bonding requirements create stringent constraints on metal selection, analogous to the selective substitutability patterns in metallic HEAs, making them suitable for testing our framework's core principle of managing uncertainty in highly selective multi-component systems. $\mathcal{D}_{0.9T_{m}}$ and $\mathcal{D}_{\text{1350K}}$: These computational datasets include \emph{all possible quaternary} alloys generated from a set of 26 elements: Fe, Co, Ir, Cu, Ni, Pt, Pd, Rh, Au, Ag, Ru, Os, Si, As, Al, Re, Mn, Ta, Ti, W, Mo, Cr, V, Hf, Nb, and Zr. The stability of these alloys is predicted using methods proposed by Chen \emph{et al.} at two different temperatures: $0.9\,T_m$ (approximately $90\%$ of the melting temperature $T_m$ of the alloy) and $1350\,( K)$. These predictions are obtained via a high-throughput computational workflow, which employs a regular-solution model using binary interaction parameters derived from \textit{ab initio} density functional theory (DFT) to compute and compare Gibbs free energies of solid solutions against competing intermetallic phases.$\mathcal{D}_{Mag}$ and $\mathcal{D}_{T_C}$: These computational datasets comprise 5,968 quaternary high-entropy alloys (HEAs), each formed by selecting four elements from a set of 21 transition metals: Fe, Co, Ir, Cu, Ni, Pt, Pd, Rh, Au, Ag, Ru, Os, Tc, Re, Mn, Ta, W, Mo, Cr, V, and Nb. Their magnetizations ($\mathcal{D}_{Mag}$) and Curie temperatures ($\mathcal{D}_{T_C}$) in the body-centered cubic (BCC) phase are computed using the Korringa--Kohn--Rostoker coherent approximation method. These datasets are derived from an original pool of $147,630$ equiatomic quaternary HEAs. $\mathcal{D}_{\text{HEA}}^{\text{exp}}$: The experimental dataset includes 55 experimentally verified quaternary HEAs from peer-reviewed publications. The dataset includes both HEA (40 alloys) and non-HEA (15 alloys) compositions, providing balanced representation for validation.$\mathcal{D}_{\text{HEB}}^{\text{exp}}$: The experimental dataset includes 19 experimentally verified quinary HEBs from peer-reviewed publications. The dataset includes 15 quinary systems forming HEB.
开发具有理想性能的新型高熵合金(high-entropy alloys, HEAs)面临着广阔成分空间与相形成机制复杂性的双重挑战。现有多款擅长插值任务的归纳筛选方法已被提出,但此类方法在外推至新型合金体系时表现欠佳。本研究提出一种框架,可通过大语言模型(large language models, LLMs)系统性整合材料数据集提取的知识与科学文献中的专家知识,从而突破外推局限。该框架的核心是元素替换原则,即识别出在保留目标性能的前提下可相互替换的化学相似元素。为建模并融合多源知识证据,我们采用了邓普斯特-谢弗(Dempster--Shafer)理论,该理论为不确定性推理提供了数学基础。我们的框架在所有实验中均优于依赖单源知识的传统相选择模型,在预测训练数据中未出现元素的成分的相稳定性时优势尤为显著。重要的是,本框架可有效补全现有方法的短板,此外还能提供可解释的推理过程,阐明对高熵合金形成过程中合金稳定性至关重要的元素替换模式。上述结果凸显了该框架在知识迁移与外推方面的潜力,为探索高熵合金广阔的成分空间提供了兼具泛化性与可解释性的高效途径。 实验阶段共考量了四类四元合金计算数据集、一类四元合金实验数据集,以及一类五元高熵硼化物(quinary high-entropy borides, HEB)实验数据集。高熵硼化物是一类单相陶瓷,多种过渡金属阳离子随机分布于硼化物结构的金属亚晶格中,兼具金属与陶瓷的独特性能组合。尽管二者键合机制存在差异,但高熵硼化物与高熵合金一样表现出极强的元素选择性:硼的严格键合要求对金属元素的选择形成了严苛限制,这与金属高熵合金中的选择性替换模式类似,因此该体系适合用于验证我们框架在高选择性多组分系统中处理不确定性的核心原则。 $mathcal{D}_{0.9T_{m}}$ 与 $mathcal{D}_{ ext{1350K}}$:这两类计算数据集涵盖了由26种元素(Fe、Co、Ir、Cu、Ni、Pt、Pd、Rh、Au、Ag、Ru、Os、Si、As、Al、Re、Mn、Ta、Ti、W、Mo、Cr、V、Hf、Nb、Zr)生成的所有可能的四元合金。采用Chen等人提出的方法,在两种温度下预测这些合金的稳定性:$0.9,T_m$(约为合金熔点$T_m$的90%)与1350K。上述预测通过高通量计算流程生成,该流程采用基于从头算密度泛函理论(density functional theory, DFT)得到的二元相互作用参数构建规则溶液模型,用于计算并对比固溶体与竞争金属间相的吉布斯自由能。 $mathcal{D}_{Mag}$ 与 $mathcal{D}_{T_C}$:此类计算数据集包含5968组四元高熵合金,每一组均从21种过渡金属(Fe、Co、Ir、Cu、Ni、Pt、Pd、Rh、Au、Ag、Ru、Os、Tc、Re、Mn、Ta、W、Mo、Cr、V、Nb)中选取四种元素组合而成。采用Korringa-Kohn-Rostoker相干近似方法,计算其体心立方(body-centered cubic, BCC)相的磁化强度($mathcal{D}_{Mag}$)与居里温度($mathcal{D}_{T_C}$)。该数据集源自147630组等原子比四元高熵合金的原始池。 $mathcal{D}_{ ext{HEA}}^{ ext{exp}}$:该实验数据集包含来自同行评议文献的55组经实验验证的四元高熵合金,其中包括40组高熵合金与15组非高熵合金成分,可实现均衡的验证样本分布。 $mathcal{D}_{ ext{HEB}}^{ ext{exp}}$:该实验数据集包含来自同行评议文献的19组经实验验证的五元高熵硼化物,其中包括15组可形成高熵硼化物的五元体系。



