CausalBench A Comprehensive Benchmark for Evaluating Causal Reasoning Capabilities of Large Language Models
收藏官方服务:
资源简介:
CausalBench is a comprehensive benchmark dataset designed to evaluate the causal reasoning capabilities of large language models. The primary uses of this dataset include, but are not limited to: - Testing the performance of large language models on causal reasoning tasks - Serving as a benchmark dataset for causal reasoning research - Improving and developing new causal reasoning algorithms and models
提供机构:
Zenodo创建时间:
2024-06-13



