BinMetric
收藏资源简介:
BinMetric是一个全面的数据集,用于评估大型语言模型在二进制分析任务上的性能。该数据集包含来自20个真实开源项目的1000个问题,涵盖了6个实际二进制分析任务,包括反编译、代码摘要、汇编指令生成等,反映了实际的逆向工程场景。数据集的设计考虑了真实二进制分析场景的复杂性,旨在提供一个标准化的评估框架,以评估大型语言模型在关键领域的有效性。
BinMetric is a comprehensive dataset designed to evaluate the performance of large language models (LLMs) on binary analysis tasks. It contains 1000 questions sourced from 20 real-world open-source projects, covering 6 practical binary analysis tasks including "decompilation", "code summarization", "assembly instruction generation", and others, which reflects real-world reverse engineering scenarios. The dataset is developed with full consideration of the complexity of actual binary analysis scenarios, aiming to provide a standardized evaluation framework for assessing the effectiveness of large language models in key domains.
BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
数据集概述
- 标题: BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
- 作者: Xiuwei Shang, Guoqiang Chen, Shaoyin Cheng, Benlong Wu, Li Hu, Gangyang Li, Weiming Zhang, Nenghai Yu
- 提交日期: 12 May 2025
- 领域: Computer Science > Software Engineering
- arXiv标识符: arXiv:2505.07360v1 [cs.SE]
- DOI: https://doi.org/10.48550/arXiv.2505.07360
数据集详情
-
摘要:
- BinMetric是一个专门用于评估大型语言模型在二进制分析任务上性能的综合基准。
- 包含1,000个问题,源自20个真实世界的开源项目,涵盖6个实用的二进制分析任务(如反编译、代码摘要、汇编指令生成等)。
- 旨在反映实际逆向工程场景,填补该领域标准化基准的空白。
- 通过实证研究揭示了当前最先进大型语言模型在二进制分析中的优势和局限性。
-
任务类型:
- 反编译
- 代码摘要
- 汇编指令生成
- 其他二进制分析任务
-
数据来源:
- 20个真实世界的开源项目
相关论文信息
- 评论: 23页,5张图,将发表于IJCAI 2025
- 引用格式: arXiv:2505.07360 [cs.SE]
访问链接
- PDF: http://arxiv.org/pdf/2505.07360v1
- HTML (实验性): http://arxiv.org/html/2505.07360v1
- TeX源码: http://arxiv.org/format/2505.07360v1




