CIRCA: Clinical Intent Representation, Cross-corpus Annotations
收藏资源简介:
CIRCA is a dataset of clinical intents annotated in the Clinical Intent Representation (CIR), harmonized across five heterogeneous source corpora (CLIP, MedDec, ap_parsing, PaniniQA, and SIMORD/ACI-Bench). Version 2 contains 19,758 CIR records (19,503 intents plus 255 human-rejected hard negatives) over 1,170 notes (1,168 with at least one intent), each with 11 content fields, including two novel axes, request_intent (clinical authority, adapted from FHIR R4 RequestIntent) and modality (clinical strength, a CIR axis with no RequestIntent equivalent). Every record carries a provenance tier: human_gold (2,145; human-reviewed and confirmed, reviewer corrections applied), human_rejected (255; reviewed, not a prospective intent), model_agreement (3,314; all three annotation models agreed), and unreviewed (14,044; models disagreed, outside the reviewed sample). This deposit publishes annotations, value mappings, benchmark files (splits, span set, span-level model predictions including a fine-tuned encoder, second-annotator labels and model predictions on them, exhaustive gold), and reconstruction and evaluation code. It contains no MIMIC-III note text other than short normalized target concept labels; the public ACI-Bench notes are included. MIMIC-III-derived annotations are released as stand-off records (character offsets plus SHA-256) and are rebuilt from a user's own PhysioNet-credentialed MIMIC-III copy using the included hydrator and KART mappings; the ACI-Bench layer (public, simulated; text CC BY 4.0, SIMORD source labels CDLA-Permissive-2.0) is complete. See README.md for the reconstruction procedure and the data-provenance and compliance statement. Version 2.2 fixes the evaluator, FHIR mapper, and hydrator robustness issues, adds a near-duplicate flag, the human-reference predictions, and the fine-tuned encoder; labels, counts, offsets, and hashes are unchanged from 2.1.



