Multidimensional Evaluation Framework for Local LLM-Based Clinical Summarization — Data Deposit
收藏资源简介:
Zenodo Deposit Multidimensional Evaluation Framework for Local LLM-Based Clinical Summarization Authors: Luis Vera, Olga Craveiro, Ricardo Malheiro, Manuel Dias, Ricardo Correia Bezerra Journal: Machine Learning and Knowledge Extraction (MDPI MAKE) DOI: 10.5281/zenodo.22962371 License: CC BY 4.0 Contents Folder What it contains 01_factuality_results/ Per-variant S/NS/C labels for all 6190 claims (case_hash + label only; clinical text stripped). Summary JSONs and Wilson CIs. 02_multijudge_agreement/ 3-model external panel consensus labels (stripped). Agreement analysis JSONs (κ matrices). Per-variant κ. 03_sensitivity_analysis/ Sensitivity by imaging modality (CT vs non-CT). Temperature experiment (R04) summary. PT→PT vs PT→EN cross-lingual comparison. 04_evaluation_schemas_and_prompts/ Expert review rubric scores (7 dimensions). Error taxonomy per-case counts. Prompt templates P1/P2/P3 with {source_text} placeholder. Experiment config example. 05_model_sha256_and_hyperparams/ Ollama blob SHA-256 digests for all 4 models. Generation and annotation hyperparameters. 06_scripts/ All Python analysis and evaluation scripts. Privacy note The underlying PACS-derived imaging reports and instantiated prompts containing real clinical text are NOT included and cannot be released (data-custodian agreement). All annotation CSVs have been stripped of the generated_claim, source_support_quote_or_note, and raw_response columns. Aggregate statistics and per-claim labels are retained. Reproducibility To reproduce the SHA-256 verification, run verify_ollama_sha.py on a machine with Ollama installed and the models pulled. Generation and annotation can be re-run with run_ollama_generation.py and annotate_factuality.py on any FHIR DiagnosticReport corpus.



