遇见数据集

llm-agents/CriticBench

收藏
Hugging Face2024-02-23 更新2024-03-04 收录
官方服务:

资源简介:

--- license: mit task_categories: - question-answering - text-classification - text-generation language: - en paperswithcode_id: criticbench pretty_name: CriticBench size_categories: - 1K<n<10K tags: - llm - reasoning - critique - correction - discrimination - math-word-problems - math-reasoning - question-answering - commonsense-reasoning - code-generation - symbolic-reasoning - algorithmic-reasoning --- # Dataset Card for Dataset Name <!-- Provide a quick summary of the dataset. --> CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families. ## Dataset Details ### Dataset Description <!-- Provide a longer summary of what this dataset is. --> - **Curated by:** THU - **Funded by [optional]:** [More Information Needed] - **Shared by [optional]:** [More Information Needed] - **Language(s) (NLP):** EN - **License:** MIT ### Dataset Sources [optional] <!-- Provide the basic links for the dataset. --> - **Repository:** https://github.com/CriticBench/CriticBench - **Paper [optional]:** https://arxiv.org/pdf/2402.14809.pdf - **Demo [optional]:** https://criticbench.github.io/ ## Uses <!-- Address questions around how the dataset is intended to be used. --> ### Direct Use <!-- This section describes suitable use cases for the dataset. --> [More Information Needed] ## Dataset Structure <!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. --> [More Information Needed] ## Dataset Creation ### Curation Rationale <!-- Motivation for the creation of this dataset. --> [More Information Needed] ### Source Data <!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). --> #### Data Collection and Processing <!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. --> [More Information Needed] #### Who are the source data producers? <!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. --> [More Information Needed] ### Annotations [optional] <!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. --> #### Annotation process <!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. --> [More Information Needed] #### Who are the annotators? <!-- This section describes the people or systems who created the annotations. --> [More Information Needed] #### Personal and Sensitive Information <!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. --> [More Information Needed] ## Bias, Risks, and Limitations <!-- This section is meant to convey both technical and sociotechnical limitations. --> [More Information Needed] ### Recommendations <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. --> Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations. ## Citation [optional] <!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. --> **BibTeX:** ``` @misc{lin2024criticbench, title={CriticBench: Benchmarking LLMs for Critique-Correct Reasoning}, author={Zicheng Lin and Zhibin Gou and Tian Liang and Ruilin Luo and Haowei Liu and Yujiu Yang}, year={2024}, eprint={2402.14809}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` **APA:** Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., & Yang, Y. (2024). CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. ## Glossary [optional] <!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. --> [More Information Needed] ## More Information [optional] [More Information Needed] ## Dataset Card Authors [optional] [More Information Needed] ## Dataset Card Contact [More Information Needed]

许可证: MIT许可证 任务类别: - 问答 - 文本分类 - 文本生成 语言: - 英语 PapersWithCode编号: criticbench 友好名称: CriticBench 样本规模类别: - 1000 < 样本量 < 10000 标签: - 大语言模型(LLM) - 推理 - 评判 - 修正 - 判别 - 数学应用题 - 数学推理 - 问答 - 常识推理 - 代码生成 - 符号推理 - 算法推理 # 数据集卡片 CriticBench是一款综合性基准测试集,旨在评估大语言模型(LLM)在多种任务场景下的生成、评判/判别与推理修正能力。CriticBench涵盖五大推理领域:数学推理、常识推理、符号推理、代码生成与算法推理。该数据集整合了15个子数据集,并纳入了三类大语言模型家族的生成结果。 ## 数据集详情 ### 数据集概述 - **整理方**:清华大学(THU) - **资助方 [可选]**:[需补充更多信息] - **共享方 [可选]**:[需补充更多信息] - **自然语言处理所用语言**:英语 - **许可证**:MIT许可证 ### 数据集来源 [可选] - **代码仓库**:https://github.com/CriticBench/CriticBench - **相关论文 [可选]**:https://arxiv.org/pdf/2402.14809.pdf - **演示页面 [可选]**:https://criticbench.github.io/ ## 使用场景 ### 直接使用 [需补充更多信息] ## 数据集结构 [需补充更多信息] ## 数据集构建 ### 构建初衷 [需补充更多信息] ### 源数据 #### 数据收集与处理流程 [需补充更多信息] #### 源数据生产者 [需补充更多信息] ### 标注信息 [可选] #### 标注流程 [需补充更多信息] #### 标注者 [需补充更多信息] #### 个人与敏感信息说明 [需补充更多信息] ## 偏差、风险与局限性 [需补充更多信息] ### 建议 用户应知晓该数据集存在的风险、偏差与局限性,相关建议仍需进一步补充完善。 ## 引用信息 [可选] **BibTeX格式引用**: @misc{lin2024criticbench, title={CriticBench: Benchmarking LLMs for Critique-Correct Reasoning}, author={Zicheng Lin and Zhibin Gou and Tian Liang and Ruilin Luo and Haowei Liu and Yujiu Yang}, year={2024}, eprint={2402.14809}, archivePrefix={arXiv}, primaryClass={cs.CL} } **APA格式引用**: Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., & Yang, Y. (2024). CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. ## 术语表 [可选] [需补充更多信息] ## 更多信息 [可选] [需补充更多信息] ## 数据集卡片撰写者 [可选] [需补充更多信息] ## 数据集卡片联系人 [需补充更多信息]

提供机构:
llm-agents
原始信息汇总

数据集卡片 for CriticBench

数据集概述

CriticBench 是一个综合基准,旨在评估大型语言模型(LLMs)在生成、批判/区分和纠正推理方面的能力,涵盖五个推理领域:数学、常识、符号、编码和算法。它整合了15个数据集,并包含了来自三个LLM家族的响应。

数据集详情

数据集描述

  • :THU 策划
  • 语言:英语(EN)
  • 许可证:MIT

数据集来源

  • 仓库:https://github.com/CriticBench/CriticBench
  • 论文:https://arxiv.org/pdf/2402.14809.pdf
  • 演示:https://criticbench.github.io/

使用

直接使用

[更多信息待补充]

数据集结构

[更多信息待补充]

数据集创建

策划理由

[更多信息待补充]

源数据

数据收集和处理

[更多信息待补充]

源数据生产者

[更多信息待补充]

标注

标注过程

[更多信息待补充]

标注者

[更多信息待补充]

个人和敏感信息

[更多信息待补充]

偏差、风险和局限性

[更多信息待补充]

建议

用户应意识到数据集的风险、偏差和局限性。更多信息待补充以供进一步建议。

引用

BibTeX:

@misc{lin2024criticbench, title={CriticBench: Benchmarking LLMs for Critique-Correct Reasoning}, author={Zicheng Lin and Zhibin Gou and Tian Liang and Ruilin Luo and Haowei Liu and Yujiu Yang}, year={2024}, eprint={2402.14809}, archivePrefix={arXiv}, primaryClass={cs.CL} }

APA:

Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., & Yang, Y. (2024). CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.

搜集汇总
数据集介绍
llm-agents/CriticBench 数据集图片
构建方式
CriticBench是一个专为评估大语言模型在推理任务中的生成、批判/辨别与修正能力而设计的综合性基准数据集。其构建过程精心整合了来自数学、常识、符号、代码及算法五大推理领域的15个公开数据集,并系统性地纳入了三个不同大语言模型家族生成的回答。通过这种多源异构数据的融合与模型响应的配对,CriticBench为深入剖析模型在复杂推理链条中的自我反思与纠错能力提供了结构化、标准化的评估框架。
特点
该数据集的核心特色在于其多维度的评估视角,不仅考察模型生成推理过程的能力,更侧重于模型对自身或他人推理进行批判性鉴别与修正的表现。涵盖的推理域广泛而全面,从严谨的数学问题到开放的常识推理,再到结构化的代码与符号逻辑,确保了评估的深度与广度。通过精心设计的任务与响应配对,CriticBench能够精细地揭示模型在推理不同阶段的薄弱环节,为理解与提升LLM的推理稳健性提供了独特洞见。
使用方法
CriticBench的使用方式灵活多样,可直接用于监督式微调与评估。研究人员可加载数据集中的问答对与模型回答,利用其标注信息训练模型提升推理与自我修正能力。在评估场景中,通过对比模型生成的推理、批判与修正结果与标准答案,可量化模型在各项子任务上的表现。该数据集兼容标准的文本生成与问答评估流程,便于集成至现有的大语言模型训练与评测管线中,推动推理与批判性思维能力的进阶研究。
背景与挑战
背景概述
CriticBench是由清华大学研究团队于2024年创建的一项综合性基准测试,旨在系统评估大语言模型(LLMs)在推理过程中的生成、批判与修正能力。随着LLMs在数学推理、常识推理、符号推理、代码生成及算法推理等五大核心领域的广泛应用,模型在复杂任务中不仅需要生成正确解答,更需具备自我反思与纠错的能力。CriticBench整合了15个公开数据集,并引入来自三个不同LLM家族的响应,构建了一个多维度、多层次的评估框架。该基准的提出填补了现有评估体系中对模型批判性思维与自我修正能力系统衡量的空白,为研究LLMs的推理鲁棒性与可解释性提供了重要工具,对推动AI在可靠性与安全性方向的发展具有深远影响。
当前挑战
CriticBench所应对的核心挑战在于LLMs在推理过程中普遍存在的自我纠错能力不足问题。传统评估多聚焦于最终答案的正确性,忽略了模型在推理链条中对错误步骤的识别与修正能力。具体挑战包括:一是构建高质量的多领域推理数据集,需确保覆盖数学、常识、代码等不同逻辑类型的任务,同时保证数据标注的准确性与多样性;二是设计有效的评估指标,以量化模型在生成、批判与修正三个阶段的性能差异,避免单一指标带来的偏见;三是处理来自不同LLM家族的响应异质性,需统一评价标准以公平比较模型能力。此外,构建过程中还需应对数据规模有限(1K-10K样本)带来的统计显著性挑战,以及如何避免过度拟合特定模型或任务类型的问题。
常用场景
经典使用场景
CriticBench作为大语言模型批判性推理能力的综合性评测基准,其经典使用场景聚焦于评估模型在数学、常识、符号、代码和算法五大推理领域中的生成、批判/判别与修正能力。研究者通过该基准可系统性地考察LLM在复杂推理链条中自我检视与纠错的表现,从而揭示模型在逻辑一致性、错误识别与修正策略上的深层局限。
解决学术问题
该数据集致力于解决大语言模型在推理任务中缺乏可靠自我评估与修正机制这一核心学术难题。传统评测多关注模型最终答案的正确性,而CriticBench通过引入批判-修正范式,使得研究者能够量化分析模型在推理过程中的错误检测灵敏度、修正有效性以及判别准确性,为理解LLM的元认知能力提供了标准化评估框架,推动了推理增强与自我反思方向的理论进展。
衍生相关工作
CriticBench的发布催生了一系列关于LLM自我反思与迭代修正的经典工作,如基于该基准提出的多轮批判-修正训练策略、强化学习驱动的推理优化方法,以及将批评信号融入模型微调过程的框架。这些衍生研究不仅深化了对模型推理机制的理解,还推动了如Self-Refine、CRITIC等代表性方法的改进,形成了以批判能力为核心的新兴研究脉络。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务