MemoryBench-Results
收藏资源简介:
本数据集是MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems(MemoryBench:LLM系统中记忆与持续学习的基准测试)的实验结果存档,旨在评估大型语言模型(LLM)系统在服务期间从累积用户反馈中学习的能力。数据集存储了已发布运行的模型预测结果、逐样本评估细节、聚合摘要以及经过清理的运行配置,便于研究社区在不重新运行完整实验的情况下审查和复用结果表格。目前包含基于四个骨干模型(Qwen3-8B、Qwen3-32B、Mistral-Small-3.2-24B-Instruct-2506和DeepSeek-V4-Flash)的结果组,每个结果组对应特定的实验配置组合,包括实验类型(如off-policy)、骨干模型、数据集分割(如domain或task)、集合名称(例如Open-Domain、Academic&Knowledge、Long-Long)以及记忆系统基线。数据以结构化目录形式组织,每个结果组目录下包含summary.json(聚合指标)、evaluate_details.json(逐样本评估细节)、predict.json(模型预测)和run_config.json(运行配置)四个文件。该数据集适用于机器学习、自然语言处理领域的研究,特别是LLM记忆机制、持续学习、基准测试评估和实验结果可复现性分析等任务。
This dataset is an experimental results archive for MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems, designed to evaluate the ability of large language model (LLM) systems to learn from accumulated user feedback during service. It stores published run results including model predictions, per-sample evaluation details, aggregated summaries, and cleaned run configurations, enabling the research community to review and reuse result tables without rerunning full experiments. Currently, it contains result groups based on four backbone models (Qwen3-8B, Qwen3-32B, Mistral-Small-3.2-24B-Instruct-2506, and DeepSeek-V4-Flash), each corresponding to specific experimental configuration combinations such as experiment type (e.g., off-policy), backbone model, dataset split (e.g., domain or task), collection names (e.g., Open-Domain, Academic&Knowledge, Long-Long), and memory system baselines. The data is organized in a structured directory format, with each result group directory containing four files: summary.json (aggregated metrics), evaluate_details.json (per-sample evaluation details), predict.json (model predictions), and run_config.json (run configuration). This dataset is suitable for research in machine learning and natural language processing, particularly for tasks related to LLM memory mechanisms, continual learning, benchmark evaluation, and experimental result reproducibility analysis.




