C-REACT: Contextualized Race and Ethnicity Annotations for Clinical Text
收藏资源简介:
The Contextualized Race and Ethnicity Annotations for Clinical Text (C-REACT) dataset is a large publicly available corpus of sentences from clinical notes manually annotated for information related to race and ethnicity (RE). The corpus presented here contains 17,281 sentences drawn from 12,000 patients and their clinical notes at the Beth Israel Deaconess Medical Center critical care units between 2001 and 2012. This corpus contains two sets of reference standard annotations for RE data. The first set contains granular RE- information such as patient country of origin and spoken language. The second set of annotations contains RE labels manually assigned by clinicians. This corpus is intended to improve understanding about granular information related to RE contained within the clinical note and how this information might be used to infer RE.



