HYPOBENCH
收藏资源简介:
HYPOBENCH是一个包含7个真实世界任务和5个合成任务的数据集集合,由芝加哥大学和多伦多大学共同构建。该数据集旨在评估大型语言模型在假设生成方面的性能,涵盖了12个领域,包括总统选举、大学入学、个性预测等。数据集结合了实际观测和现有文献,以评估假设生成的实用性和普遍性。通过控制合成数据集的难度,可以对模型发现真实假设的能力进行精确评估,从而为改进AI系统在科学研究中的应用提供有价值的信息。
HYPOBENCH is a dataset collection comprising 7 real-world tasks and 5 synthetic tasks, co-developed by the University of Chicago and the University of Toronto. This dataset is designed to evaluate the performance of large language models (LLMs) in hypothesis generation, covering 12 domains such as presidential elections, college admissions, personality prediction, and more. It integrates real-world observations and existing literature to assess the practicality and generalizability of hypothesis generation. By controlling the difficulty level of the synthetic datasets, precise evaluation of models' capability to discover genuine hypotheses can be conducted, thus providing valuable insights for improving the application of AI systems in scientific research.

- 1HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation芝加哥大学,多伦多大学 · 2025年



