遇见数据集

HILI LLM causality-evidence extraction benchmark dataset

收藏
Zenodo2026-06-19 更新2026-06-28 收录
官方服务:

资源简介:

This dataset accompanies the manuscript “Benchmarking Large Language Models for Extracting Causality-Relevant Evidence From Liver Injury Case Narratives.” The repository contains derived data products for a benchmark evaluating large language models (LLMs) as structured evidence-extraction aids for liver injury case narratives. The benchmark includes 120 held-out LiverTox-derived case narratives, 40 predefined causality-relevant evidence fields per case, evidence-span anchored reference annotations, raw model-output records, field-level scoring outputs, adjudication products, final manuscript tables, data-derived figures, prompt/schema materials, and validation/scoring scripts. The benchmark evaluates schema adherence, field-level extraction reliability, evidence-span grounding, and safety-relevant extraction failures. It does not evaluate automated RUCAM scoring, clinical diagnosis, or autonomous herb-induced or drug-induced liver injury causality assessment. Full source case narratives and populated per-case prompt files are not redistributed because source narratives may be subject to copyright or source-specific terms. The repository provides source identifiers, source URLs where available, derived annotations, short evidence spans for traceability, model outputs, scoring products, and scripts to support auditability and reproducibility.

提供机构:
Zenodo
创建时间:
2026-06-19
二维码
社区交流群
二维码
科研交流群
商业服务