APT Knowledge Graph Dataset (Embedded Data)
收藏资源简介:
This repository provides embedded representations of the APT knowledge graph dataset, as described in Zhou et al., 2025. To comply with confidentiality requirements, only processed embeddings are shared, enabling researchers to conduct threat analysis and modeling without accessing sensitive raw data. Dataset Overview The explicit relation dataset uses the APT knowledge graph with 10 entity types and 11 relationship types. Data statistics are presented in Table 1: Node Type Quantities and Table 2: Relationship Type Quantities. For implicit relationships, particularly the (Indicator → Uses → TTPs) triples, the dataset includes: 10,127 Indicators: Malicious samples and analysis results, representing mutable "symptom"-level features in the "Pyramid of Pain" model. 293 TTPs: Attacker behaviors and strategies at the apex of the "Pyramid of Pain." 148,231 triples: Sourced from VirusTotal, providing a reliable benchmark for inferring Indicator-TTP relationships. The embedded data in this repository preserves the structural and semantic properties of the original dataset while ensuring confidentiality. Node Types The nodes.csv file details the following node types: Node Type ID Node Type 1 Attacker 2 TTPs 3 Campaign 4 CourseOfAction 5 Identity 6 Indicator 7 Location 8 Malware 9 Tool 10 Vulnerability Dataset Statistics Table 1: Node Type Quantities Node Type Quantity Attacker 149 TTPs 883 Campaign 76 CourseOfAction 57 Identity 305 Indicator 19,287 Location 250 Malware 1,174 Tool 268 Vulnerability 296,978 Total 319,427 Table 2: Relationship Type Quantities Relationship Type Quantity Attributed to 95 Authored by 54 Downloads 53 Exploits 518 Impersonates 52 Indicates 26,969 Located at 582 Mitigates 1,608 Related to 23 Targets 1,141 Uses 15,996 Total 47,091 Why Embedded Data? The original APT knowledge graph contains sensitive information unsuitable for public release. The provided embeddings retain essential structural and semantic properties, making them ideal for research tasks such as: Inferring implicit Indicator-TTP relationships. Mapping low-level attack traces to high-level behavioral patterns. Enhancing threat detection and traceability. Data Source and Rationale The implicit (Indicator → Uses → TTPs) triples were sourced from VirusTotal for the following reasons: The APT knowledge graph lacks explicit (Indicator, Uses, TTPs) triples. Indicators represent mutable features at the base of the "Pyramid of Pain," while TTPs reflect high-level attacker behaviors. Inferring these relationships improves threat detection and traceability. VirusTotal's public API provides reliable TTP information linked to Indicators, serving as a robust benchmark for validating inferred relationships. Usage The embedded data is suitable for research applications, including training machine learning models for cybersecurity threat analysis and evaluating knowledge graph completion methods. Refer to the provided data files and documentation for details on the embedding format and usage instructions. Contact For questions or additional information, please contact th



