HiDy: A Large-scale Hierarchical Dynamic Financial Knowledge Base
收藏资源简介:
In recent years, domain-specific knowledge bases (KBs) have been increasingly popular due to their expertise and in-depth representation in a specific domain. Among these domain-specific KBs, financial KBs have attracted considerable attention from academics and industries due to their broad spectrum of downstream applications, such as stock movement prediction, financial fraud detection, and supply chain management. Until now, no datasets including open-domain and domain-specific KBs provide dynamic, up-to-date, and comprehensive financial knowledge, which leads to a fair comparison and broader research among popular financial task models impossible or at least very difficult. We, therefore, construct a dynamic financial KB, HiDy, to address the current limitations discussed above. HiDy is a hierarchical, dynamic, robust, diverse, and large-scale financial KB that aims to provide various valuable financial knowledge as critical benchmarking data for fair model testing in different financial tasks. Specifically, HiDy currently contains 34 relation types, more than 492,600 relations, 17 entity types, and more than 51,000 entities. The scale of HiDy is steadily growing due to its continuous updates. To make HiDy easily accessible and retrieved, HiDy is organized in a well-formed financial hierarchy with four branches, <em>Macro</em>, <em>Meso</em>,<em> Micro</em>, and<em> Others</em>. Moreover, the robustness of HiDy is a result of the state-of-the-art knowledge extraction and fusion techniques as well as the manual cleaning. For temporality, HiDy's dynamic knowledge is extracted from various resources under different time-granularity to ensure each branch's knowledge is the most up-to-date.
近年来,领域特定知识库(domain-specific knowledge bases,KBs)因能够在特定领域内实现专业的知识储备与深度表征而愈发受到青睐。在此类领域特定知识库中,金融知识库(financial KBs)凭借其广泛的下游应用场景,例如股价走势预测、金融欺诈检测以及供应链管理等,吸引了学界与产业界的广泛关注。截至目前,尚无同时涵盖开放域与领域特定知识库的数据集可提供动态、实时且全面的金融知识,这使得针对主流金融任务模型开展公平对比与更广泛的研究变得难以实现,即便可行,难度也极高。为此,我们构建了动态金融知识库HiDy,以解决上述现存的局限性问题。HiDy是一款具备层级结构、动态特性、鲁棒性、多样性与大规模特性的金融知识库,旨在提供丰富且极具价值的金融知识,作为关键的基准数据集,用于各类金融任务中的模型公平测试。具体而言,HiDy当前包含34种关系类型、逾492600条关系、17种实体类型以及超过51000个实体。得益于持续的更新维护,HiDy的规模正稳步扩张。为便于HiDy的获取与检索,该知识库以结构严谨的金融层级体系进行组织,包含四大分支:<em>宏观(Macro)</em>、<em>中观(Meso)</em>、<em>微观(Micro)</em>与<em>其他(Others)</em>。此外,HiDy的鲁棒性源于采用了当前最先进的知识抽取与融合技术,同时辅以人工清洗流程。在时效性方面,HiDy的动态知识源自不同时间粒度的多源数据,以确保各分支下的知识始终保持最新状态。



