Data-Agent-Evaluation-Dataset
收藏资源简介:
# Data-Agent-Evaluation-Dataset ## Dataset Description **Data-Agent-Evaluation-Dataset** is an evaluation dataset designed for domain-specific large language models (LLMs), covering three key vertical domains: **finance, medicine, and law**. It integrates original data and processed evaluation samples from multiple authoritative benchmarks, providing a standardized and reproducible foundation for assessing models’ knowledge mastery capabilities. This dataset is intended to be used with the [Data-Agent-Evaluation](https://github.com/haolpku/Data-Preparation-Bench) evaluation platform, supporting automated evaluation processes to ensure fairness and scientific rigor. ## Data Sources and Composition The dataset consists of the following subsets: | Domain | Included Benchmarks | Description | |----------|---------------------|-------------| | Finance | FinCDM, XFinBench | Covers financial reading comprehension, numerical reasoning, compliance analysis, etc. | | Medicine | MedCaseReasoning, MedmcQA, MedRBench | Includes medical case reasoning, professional Q&A, clinical decision support, etc. | | Law | legalbench, lex-glue | Contains legal text understanding, case classification, contract analysis, etc. | Each subset retains the original task definitions. Overly subjective or open-ended questions have been filtered out to ensure the evaluation focuses on knowledge acquisition and logical reasoning. ## Usage Restrictions and License This dataset is released under the respective licenses of the original benchmarks. Users must comply with the original terms of use for each benchmark and properly credit the data sources. The data is intended for academic research and non-commercial use only. Any illegal or unethical use is prohibited.



