遇见数据集

FHIRBench: Benchmark Data and Evaluation Results

收藏
Zenodo2026-07-06 更新2026-08-02 收录
官方服务:

资源简介:

Raw evaluation data for FHIRBench: Benchmarking FHIR Clinical Data Serialization Strategies for Large Language Models. Contains Layer 1 (token-level F1) scored responses for 4 models x 1800 evaluations each, Layer 2 (LLM-as-judge) scores for 7187 evaluations, patient-level statistical tests, and the canonical prompt sets used for evaluation.

提供机构:
Zenodo
创建时间:
2026-07-06
二维码
社区交流群
二维码
科研交流群
商业服务