BGSynthMedic: Synthetic Parallel English-Bulgarian Discharge Summary Corpus with Automatic Annotations
收藏资源简介:
Data scarcity is a significant problem for advancing clinical Natural Language Processing (NLP) in low‑resource languages like Bulgarian. To address this issue, we present BGSynthMedic, a publicly available synthetic parallel corpus of English–Bulgarian discharge summaries with automatic annotations for five entity types and their Unified Medical Language System (UMLS) concept links. We use the SynthMedic dataset, a corpus of 450 synthetic discharge summaries generated using medical guidelines and GPT-4, and automatically translate it into Bulgarian with BgGPT. We use zero-shot multilingual and monolingual Named Entity Recognition (NER) models from the GLiNER family to automatically label diseases, symptoms, procedures, medication, and observations in the original and translated summaries. The extracted entities are automatically linked to the UMLS. The machine translation quality is evaluated using reference-free metrics, yielding a BERTScore of 0.90 and a LaBSE of 0.89. The presented pipeline demonstrates how to create a parallel corpus of synthetic discharge summaries and a set of silver-standard annotations in a low-resource language. We release the BGSynthMedic dataset, including parallel medical term dictionaries and abbreviations, as a community resource to facilitate future research in Bulgarian clinical NLP.



