遇见数据集

soap-sbar-llm-benchmark: Fact-Transcription Fidelity in LLM SOAP-to-SBAR Translation (Dataset and Pipeline), v1.1.0

收藏
Zenodo2026-05-10 更新2026-05-26 收录
官方服务:

资源简介:

Reproducibility package for the three-stage automated evaluation pipeline applied to fact-transcription fidelity benchmarking of three commercial large language models (GPT-4o, Claude Opus 4.5, Gemini 2.5 Flash-Lite) on SOAP-to-SBAR translation in nursing handoff documentation. Contents 9 source SOAP scenarios, authored de novo across acute, chronic, and palliative-care categories (3 per category). 180 pre-specified fact tags with annotations (content domain, clinical priority, target SBAR section, preservation difficulty stratum). 1,620 generated SBAR records (3 vendors × 9 scenarios × 2 sampling temperatures × 30 trials). Three-stage automated evaluation pipeline source code (deterministic Stage 1 + paired AI judges Stage 2 + arbiter Stage 3). Statistical analysis scripts (Friedman, Wilcoxon signed-rank, Cohen's κ, ICC(2,1), Kruskal-Wallis). Per-call resolved_model audit logs. Pinned dependency versions in requirements.txt. Changes in v1.1.0 This release synchronises the repository with the post-correction manuscript submitted to Cureus. No changes to raw data, generated SBAR records, or analysis-pipeline source code; corrections are limited to documentation, manuscript-aligned metadata, and the rendering of Figure 4. See CHANGELOG.md in the repository for full details. Linked manuscript Tajima H. Fact-Transcription Fidelity in Large Language Model SOAP-to-SBAR Translation: A Three-Stage Automated Evaluation Across Three Vendors. Cureus, 2026; in press. Licence MIT Licence. See LICENSE in the repository.

提供机构:
Zenodo
创建时间:
2026-05-10
二维码
社区交流群
二维码
科研交流群
商业服务