BLADE
收藏资源简介:
BLADE是一个模块化和可扩展的基准测试框架,旨在评估由大型语言模型(LLM)驱动的自动算法发现(AAD)方法。该框架集成了多个基准问题(包括MA-BBOB和SBOX-COST等)的集合,旨在进行能力导向的测试,例如泛化、专业化和信息利用。BLADE提供了灵活的实验设置选项,标准化日志记录以确保可重复性和公平比较,并包含用于分析AAD过程的方法(例如代码演化图和多种可视化方法),并通过与IOHanalyser和IOHexplainer等现有工具的集成,便于与人工设计的基线进行比较。
BLADE is a modular and extensible benchmarking framework designed to evaluate automated algorithm discovery (AAD) methods powered by large language models (LLMs). It integrates a suite of benchmark problems including MA-BBOB, SBOX-COST, among others, to support capability-oriented testing such as generalization, specialization, and information utilization. BLADE provides flexible experimental setup options, standardized logging to ensure reproducibility and fair comparison, and includes methods for analyzing the AAD process—for example, code evolution graphs and various visualization techniques. Additionally, it facilitates comparison with manually designed baselines through integration with existing tools such as IOHanalyser and IOHexplainer.




