A bilingual Italian and English dataset of clinical case questions from medical specialization exams
收藏资源简介:
ITAMed is a bilingual Italian and English dataset of 1,260 multiple-choice questions from the entrance examinations to Italian medical specialization schools (Concorso per l'accesso alle Scuole di Specializzazione in Medicina), covering nine consecutive years from 2017 to 2025. Each question has five answer options and a single correct option, and is provided in the original Italian and in an expert-reviewed English translation. Questions were extracted from official examination documents, verified against the source, and annotated by specialty using a dual-model protocol (Claude Opus 4.8 and GPT-5.5) with independent expert adjudication by two physicians, yielding one or two of 28 specialty categories per item. Seventy-six image-bearing questions (6.0%) were extracted and organised into 21 image categories through joint expert review. The data are released in XLSX and JSON formats, together with the source documents, the 76 extracted images, the processing code, and the annotation prompts, to support evaluation of language models, machine translation, medical text classification, multimodal reasoning, and medical education. Repository structure and full documentation are provided in the README.md file included in the deposit and in the following repositories: Github: https://github.com/LM-Healthcare/ITAMed Hugging Face: https://huggingface.co/datasets/Filo-White/ITAMed Contacts: Filippo Bianchini - bianchini@diag.uniroma1.it Edoardo Bianchini - e.bianchini@policlinicocampus.it



