Large-language models (LLMs) show promise for extracting information from clinical notes. Deploying these models at scale can be challenging due to high computational costs, regulatory constraints, an
本数据集包含CL4Health2026会议CRF填充共享任务的标注训练表格。临床笔记收集自意大利都灵San Giovanni Bosco医院,经过匿名化和标注处理。数据集包含两个语言分片:英语(en)和意大利语(it),共10个样本。每个样本包含:临床笔记标识符(document_id)、患者临床病史记录(clinical_note)以及带有真实标签的CRF项目标注(annotations)。
We present the SynSUM benchmark, a synthetic dataset linking unstructured clinical notes to structured background variables. The dataset consists of 10,000 artificial patient records containing t