Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery
收藏资源简介:
Discovering novel high-entropy alloys (HEAs) with desirable properties is challenging due to the vast compositional space and the complexity of phase formation mechanisms. Efficiently exploring this space requires a strategic approach that integrates diverse knowledge sources. This study proposes a framework that systematically combines knowledge extracted from computational material datasets with domain knowledge distilled from scientific literature using large language models. A central feature includes explicitly considering element substitutability and identifying chemically similar elements that can be potentially interchanged to stabilize desired HEAs. Dempster--Shafer theory, a mathematical framework for reasoning under uncertainty, is employed to model and integrate substitutability based on aggregated evidence from multiple sources. The framework predicts the phase stability of candidate HEA compositions and is systematically evaluated on quaternary alloy systems, demonstrating superior performance over baseline machine learning models and methods that rely on single-source evidence in cross-validation experiments. By leveraging multi-source knowledge, the framework retains strong predictive power even when key elements are absent from the training data, underscoring its potential for knowledge transfer and extrapolation. Furthermore, the enhanced interpretability of the methodology unveils fundamental factors governing HEA formation. Overall, this study presents a promising strategy for accelerating HEA discovery by integrating computational and textual knowledge sources, enabling efficient exploration of vast compositional spaces with improved generalization and interpretability. With the established methodological foundation, we apply the proposed method to quaternary alloy datasets, evaluating its predictive performance in terms of accuracy and interpretability. Experiments are conducted considering four computational datasets of quaternary alloys: $\mathcal{D}_{0.9Tm}$ and $\mathcal{D}_{1350K}$: These datasets include \emph{all possible quaternary} alloys generated from a set of 26 elements: Fe, Co, Ir, Cu, Ni, Pt, Pd, Rh, Au, Ag, Ru, Os, Si, As, Al, Re, Mn, Ta, Ti, W, Mo, Cr, V, Hf, Nb, and Zr. The stability of these alloys---defined as whether they form an HEA phase---is predicted using methods proposed by Chen \emph{et al.} at two different temperatures: $0.9\,T_m$ (approximately $90\%$ of the melting temperature $T_m$ of the alloy) and $1350\,( K)$. These predictions are obtained via a high-throughput computational workflow, which employs a regular-solution model using binary interaction parameters derived from \textit{ab initio} density functional theory (DFT) to compute and compare Gibbs free energies of solid solutions against competing intermetallic phases. $\mathcal{D}_{Mag}$ and $\mathcal{D}_{T_C}$: These datasets comprise 5,968 equiatomic quaternary high-entropy alloys (HEAs), each formed by selecting four elements from a set of 21 transition metals: Fe, Co, Ir, Cu, Ni, Pt, Pd, Rh, Au, Ag, Ru, Os, Tc, Re, Mn, Ta, W, Mo, Cr, V, and Nb. Their magnetizations ($\mathcal{D}_{Mag}$) and Curie temperatures ($\mathcal{D}_{T_C}$) in the body-centered cubic (BCC) phase are computed using the Korringa--Kohn--Rostoker coherent approximation method. These datasets are derived from an original pool of $147,630$ equiatomic quaternary HEAs.
开发具备优异性能的新型高熵合金(high-entropy alloys, HEAs)颇具挑战,这是因为其成分空间极其庞大,且相形成机制极为复杂。高效探索该成分空间需要整合多源知识的系统性策略。本研究提出一种框架,该框架通过大语言模型(Large Language Model)系统性整合了从计算材料数据集中提取的知识,以及从科学文献中提炼的领域知识。该框架的核心特征之一是显式考量元素可替代性,并识别化学性质相似、可潜在互换以稳定目标HEA的元素。本研究采用德普斯特-谢弗理论(Dempster-Shafer theory)——一种用于不确定性推理的数学框架——基于多源聚合证据对元素可替代性进行建模与整合。该框架可预测候选HEA组分的相稳定性,并在四元合金体系中开展系统性评估。交叉验证实验结果表明,其性能优于基线机器学习模型以及仅依赖单源证据的方法。通过利用多源知识,即便训练数据中缺失关键元素,该框架仍可保持优异的预测性能,凸显了其在知识迁移与外推方面的应用潜力。此外,该方法的可解释性得到增强,可揭示调控HEA形成的核心机制。综上,本研究提出了一种极具前景的策略,通过整合计算与文本知识源加速HEA的开发,实现对庞大成分空间的高效探索,并提升模型的泛化能力与可解释性。 在建立方法学基础后,我们将所提方法应用于四元合金数据集,并从准确性与可解释性两方面评估其预测性能。本实验基于四组四元合金计算数据集展开: $mathcal{D}_{0.9Tm}$与$mathcal{D}_{1350K}$:该数据集包含由26种元素生成的所有可能四元合金,涉及元素包括Fe、Co、Ir、Cu、Ni、Pt、Pd、Rh、Au、Ag、Ru、Os、Si、As、Al、Re、Mn、Ta、Ti、W、Mo、Cr、V、Hf、Nb及Zr。我们采用Chen等人提出的方法,在两种不同温度下预测这些合金的相稳定性(即是否形成HEA相):$0.9,T_m$(约为合金熔点$T_m$的90%)与1350K。该预测通过高通量计算流程完成:流程采用基于从头算密度泛函理论(ab initio density functional theory, DFT)得到的二元相互作用参数构建正规溶液模型,计算固溶体的吉布斯自由能,并与竞争相的金属间化合物相的吉布斯自由能进行对比。 $mathcal{D}_{Mag}$与$mathcal{D}_{T_C}$:该数据集包含5968种等原子比四元高熵合金(HEAs),这些合金均从21种过渡金属中选取4种元素组合而成,涉及元素包括Fe、Co、Ir、Cu、Ni、Pt、Pd、Rh、Au、Ag、Ru、Os、Tc、Re、Mn、Ta、W、Mo、Cr、V及Nb。我们采用Korringa-Kohn-Rostoker相干近似法(Korringa-Kohn-Rostoker coherent approximation method),计算这些合金在体心立方(body-centered cubic, BCC)相中的磁化强度(对应$mathcal{D}_{Mag}$数据集)与居里温度(对应$mathcal{D}_{T_C}$数据集)。本数据集源自147630种等原子比四元HEAs的原始候选池。



