遇见数据集

Representation Wins on QA, Not on ML: A Paired-Data Comparison of Structured FHIR and LLM-Narrative Retrieval for Adherence-Indicator Question Answering on Long-Acting Specialty Regimens

收藏
Zenodo2026-05-19 更新2026-05-26 收录
官方服务:

资源简介:

This combined deposit accompanies the preprint "Representation Wins on QA, Not on ML: A Paired-Data Comparison of Structured FHIR and LLM-Narrative Retrieval for Adherence-Indicator Question Answering on Long-Acting Specialty Regimens" (Jani 2026). The deposit contains the preprint PDF, the full source code, and the synthetic dataset, all under one DOI. Three retrieval systems are compared on adherence-indicator question answering for specialty medications: System A (Narrative RAG) over LLM-generated patient narratives; System B (Structured Naive FHIR RAG) over flat-serialised FHIR resources; and System C (Structured Resource-Aware FHIR RAG) with typed resource filtering, reference-chain traversal, and temporal pre-filtering. Two further arms isolate attribution questions: System N (no-retrieval baseline, empty context) and System A-T (deterministic templated-narrative ablation, no LLM in the narrative-generation step). Dataset: 200 synthetic patients, each rendered as a FHIR R4B bundle (validated against an R4B conformance validator), a LLM-generated narrative (Llama 3.3-70B via Groq), and a deterministic templated narrative. Both narrative sources pass an independent prompted-entity fidelity audit at 100% weighted recall. The dataset includes 13,800 programmatically-verified questions across five Patient Support Program (PSP) adherence-indicator families: next-dose lookup, dose-history aggregation (MPR numerator), coverage-window reasoning (PDC), missed-dose detection, and persistence / cross-resource component questions. All ground-truth values are computed by pure-Python adherence-metric functions over FHIR resources; no LLM-as-judge is used anywhere in evaluation. Headline result: System A achieves 40.6% exact-match accuracy on the 13,800-question evaluation, outperforming Systems B (35.3%) and C (33.4%), all pairwise differences statistically significant. The templated-narrative ablation matches and slightly exceeds the LLM-narrative arm (41.3% vs 40.6%, paired-bootstrap delta +0.74 pp [+0.13, +1.34]), demonstrating that narrative format - not LLM-specific lexical regularity - is the load-bearing property. A no-retrieval baseline (25.2% overall) decomposes the absolute accuracy into a question-text-attributable floor plus retrieval-attributable lifts of +15.4 / +10.1 / +8.2 pp for A / B / C. On a synthetic adherence classification task, FHIR-structured features substantially outperform narrative-derived and retrieval-trace features (LightGBM class-weighted AUC 0.997 vs 0.846 vs 0.769), demonstrating a representation dissociation between QA and downstream ML. No PHI: all patient data is fully synthetic. The dataset is suitable for unrestricted academic and commercial reuse under CC-BY-4.0. Three files in the deposit: - fhir_rag_preprint.pdf: 26-page preprint - fhir_rag_code_v0.1.2.zip (1.04 MB; SHA-256 b822642131914994e6dd509641bbfaeca848556423ae16fa0635278c6eb78ab4): source code, evaluation harness, analysis scripts, paper LaTeX. Code files are released under MIT (see LICENSE inside the zip); MIT is compatible with the deposit-level CC-BY-4.0 umbrella. - fhir_rag_dataset_v1.zip (52.89 MB; SHA-256 000c27a42d4a304d71e59b700e7209fce7edccbbbae7458a1671e8ffe7ce421d): 200 FHIR R4B bundles, 200 LLM narratives, 200 templated narratives, fidelity-audit reports, 13,800-question bank with programmatic ground truth, and scored output CSVs from all five system arms (A, B, C, N, A-T).

提供机构:
Zenodo
创建时间:
2026-05-18
二维码
社区交流群
二维码
科研交流群
商业服务