Bills
收藏资源简介:
Bills数据集是一个标准的主题模型基准,由马里兰大学帕克分校的研究人员创建。该数据集包含了来自第110-114届美国国会的32661个法案摘要,这些法案摘要被分类到21个顶级主题和112个次级主题中。研究通过对两个数据集的使用,评估了传统主题模型和大型语言模型在帮助用户理解大型文档集合方面的有效性,探讨了人类在环对模型性能的影响。该数据集旨在解决政策分析和理解的问题,帮助研究人员更好地探索和理解法案内容。
The Bills dataset is a standard topic model benchmark developed by researchers from the University of Maryland, College Park. It contains 32,661 bill summaries from the 110th to 114th sessions of the United States Congress, which are classified into 21 top-level topics and 112 sub-topics. The study employed two datasets to evaluate the efficacy of both traditional topic models and large language models in assisting users to comprehend large-scale document collections, and explored the impact of human-in-the-loop on model performance. This dataset aims to address challenges in policy analysis and understanding, enabling researchers to better explore and grasp the content of congressional bills.

- 1Large Language Models Struggle to Describe the Haystack without Human Help: Human-in-the-loop Evaluation of LLMs马里兰大学帕克分校 · 2025年



