AIReg-Bench
收藏资源简介:
AIReg-Bench数据集是首个用于测试大型语言模型(LLMs)在评估人工智能系统是否符合欧盟人工智能法案(AIA)方面的能力的基准数据集。该数据集包含120个技术文档摘录,每个摘录描述了一个虚构但合理的AI系统,这些摘录由LLM生成,并由法律专家标注。数据集的创建旨在提供一个基准,用于理解和评估基于LLM的AI合规性评估工具的机会和局限性,并作为后续LLMs比较的基准。
The AIReg-Bench dataset is the first benchmark dataset for testing the capabilities of large language models (LLMs) in evaluating whether artificial intelligence systems comply with the European Union's Artificial Intelligence Act (AIA). This dataset includes 120 technical document excerpts, each describing a fictional but plausible AI system. These excerpts were generated by LLMs and annotated by legal experts. The dataset was developed to provide a benchmark for understanding and evaluating the opportunities and limitations of LLM-based AI compliance assessment tools, and to serve as a baseline for subsequent comparisons of LLMs.




