HELIX
收藏资源简介:
HELIX数据集是由麻省理工学院林肯实验室开发的,用于程序相似性研究的合成数据集。该数据集通过程序切片技术从开源库中自动提取了28,178个独特的组件,这些组件代表了特定的程序功能。数据集的创建旨在解决现有数据集在程序相似性评估中的不足,特别是缺乏高质量和相关性的问题。HELIX数据集支持多种编程语言和构建系统,适用于机器学习在程序分析领域的应用,特别是在恶意软件分析和安全领域。
The HELIX dataset is a synthetic dataset developed by MIT Lincoln Laboratory for program similarity research. It automatically extracts 28,178 unique components representing specific program functionalities from open-source libraries via program slicing techniques. This dataset was developed to address the shortcomings of existing datasets in program similarity evaluation, particularly the lack of high-quality and relevant resources. The HELIX dataset supports multiple programming languages and build systems, making it suitable for machine learning applications in program analysis, especially in malware analysis and cybersecurity.




