遇见数据集

Data-Agent-Evaluation-Dataset

收藏
魔搭社区2026-07-06 更新2026-07-15 收录
官方服务:

资源简介:

# Data-Agent-Evaluation-Dataset ## Dataset Description **Data-Agent-Evaluation-Dataset** is an evaluation dataset designed for domain-specific large language models (LLMs), covering three key vertical domains: **finance, medicine, and law**. It integrates original data and processed evaluation samples from multiple authoritative benchmarks, providing a standardized and reproducible foundation for assessing models’ knowledge mastery capabilities. This dataset is intended to be used with the [Data-Agent-Evaluation](https://github.com/haolpku/Data-Preparation-Bench) evaluation platform, supporting automated evaluation processes to ensure fairness and scientific rigor. ## Data Sources and Composition The dataset consists of the following subsets: | Domain | Included Benchmarks | Description | |----------|---------------------|-------------| | Finance | FinCDM, XFinBench | Covers financial reading comprehension, numerical reasoning, compliance analysis, etc. | | Medicine | MedCaseReasoning, MedmcQA, MedRBench | Includes medical case reasoning, professional Q&A, clinical decision support, etc. | | Law | legalbench, lex-glue | Contains legal text understanding, case classification, contract analysis, etc. | Each subset retains the original task definitions. Overly subjective or open-ended questions have been filtered out to ensure the evaluation focuses on knowledge acquisition and logical reasoning. ## Usage Restrictions and License This dataset is released under the respective licenses of the original benchmarks. Users must comply with the original terms of use for each benchmark and properly credit the data sources. The data is intended for academic research and non-commercial use only. Any illegal or unethical use is prohibited.

提供机构:
maas
创建时间:
2026-04-01
二维码
社区交流群
二维码
科研交流群
商业服务