遇见数据集

External utilities of items from Table 1.

收藏
Figshare2025-02-03 更新2026-04-28 收录
官方服务:

资源简介:

Privacy is as a critical issue in the age of data. Organizations and corporations who publicly share their data always have a major concern that their sensitive information may be leaked or extracted by rivals or attackers using data miners. High-utility itemset mining (HUIM) is an extension to frequent itemset mining (FIM) which deals with business data in the form of transaction databases, data that is also in danger of being stolen. To deal with this, a number of privacy-preserving data mining (PPDM) techniques have been introduced. An important topic in PPDM in the recent years is privacy-preserving utility mining (PPUM). The goal of PPUM is to protect the sensitive information, such as sensitive high-utility itemsets, in transaction databases, and make them undiscoverable for data mining techniques. However, available PPUM methods do not consider the generalization of items in databases (categories, classes, groups, etc.). These algorithms only consider the items at a specialized level, leaving the item combinations at a higher level vulnerable to attacks. The insights gained from higher abstraction levels are somewhat more valuable than those from lower levels since they contain the outlines of the data. To address this issue, this work suggests two PPUM algorithms, namely MLHProtector and FMLHProtector, to operate at all abstraction levels in a transaction database to protect them from data mining algorithms. Empirical experiments showed that both algorithms successfully protect the itemsets from being compromised by attackers.

数据时代下,隐私保护已是至关重要的议题。各类组织机构与企业公开共享其数据时,往往面临一项核心关切:其敏感信息可能被竞争对手或攻击者借助数据挖掘工具泄露或窃取。高效用项集挖掘(High-utility itemset mining, HUIM)是频繁项集挖掘(Frequent itemset mining, FIM)的扩展分支,旨在处理以事务数据库形式存储的商业数据,而这类数据同样面临被盗取的风险。为应对这一问题,诸多隐私保护数据挖掘(privacy-preserving data mining, PPDM)技术被相继提出。近年来,隐私保护效用挖掘(privacy-preserving utility mining, PPUM)成为PPDM领域的重要研究课题,其目标是保护事务数据库中的敏感信息(如敏感高效用项集),使其无法被数据挖掘技术所发现。然而,现有PPUM方法并未考虑数据库中项的泛化机制(如类别、分类、分组等),此类算法仅针对特定粒度下的项进行处理,使得更高抽象层级的项组合极易遭受攻击。相较于低抽象层级的信息,更高抽象层级所蕴含的数据轮廓往往具备更高的价值。为解决这一问题,本研究提出两种PPUM算法:MLHProtector与FMLHProtector,可在事务数据库的所有抽象层级上运行,从而抵御各类数据挖掘算法的攻击。实证实验结果表明,这两种算法均能有效保护项集免受攻击者的破解。

创建时间:
2025-02-03
二维码
社区交流群
二维码
科研交流群
商业服务