A bilingual Italian–English dataset of case-based and knowledge-based questions from medical specialization examinations
收藏资源简介:
ITAMed is a bilingual Italian–English dataset of 1,260 multiple-choice questions from the Italian national competitive examinations for admission to medical specialization schools (Concorso per l’accesso alle Scuole di Specializzazione in Medicina, SSM). The dataset covers nine consecutive examination years, from 2017 to 2025, with 140 questions per year. Each record contains five answer options and one official keyed answer and is provided in both the original Italian and an expert-reviewed English translation. The canonical files preserve the option order of the official examination releases, in which the keyed answer is always presented in position A. A reproducible answer-shuffling utility is included for model evaluation and benchmarking. Questions were extracted from the publicly released official examination documents and manually verified against the source by two Italian-speaking physicians. For the scenario-based examinations administered between 2017 and 2019, the relevant shared clinical scenario was prepended to each associated question so that every record is self-contained. Each question is enriched with: Medical specialty annotation: one or two labels selected from a taxonomy of 28 specialty categories. Initial classification was performed independently by Claude Opus 4.8 and GPT-5.5 using the same Italian-language prompt. All 380 questions for which the two models disagreed on any component of the assignment were independently adjudicated by two physicians, with residual disagreements resolved by consensus. Question-type annotation: each item was jointly classified by two physicians as either case-based or knowledge-based, according to predefined operational criteria. The final dataset contains 928 case-based questions (73.7%) and 332 knowledge-based questions (26.3%). Image annotation: 76 questions (6.0%) contain an image extracted from the corresponding official examination document. Images were manually verified, linked to the relevant records through relative file paths, and assigned by joint expert review to one of 21 controlled image categories. Deposit contents The archive includes: complete and per-year datasets in XLSX and JSON formats; aligned Italian and English versions; specialty, question-type, and image annotations; 76 extracted image files organised by examination year; specialty, question-type, and image-distribution workbooks; the nine source examination documents; PDF extraction and image-processing scripts; specialty-classification prompts, scripts, raw model outputs, agreement analyses, independent expert-review files, and consensus records; translation prompts, scripts, intermediate outputs, correction logs, and tracked-review files; an answer-option shuffling utility for generating reproducible evaluation copies; scripts and outputs used to generate the distribution charts; software requirements and detailed README documentation for the complete reproducibility workflow. The resource is intended to support research on: medical question answering; evaluation of large and small language models; bilingual and cross-lingual benchmarking; medical machine translation; specialty and question-type classification; multimodal medical reasoning; training-data contamination and temporal-generalisation studies; retrieval-augmented generation; medical education. Additional repositories GitHub:https://github.com/LM-Healthcare/ITAMed Hugging Face:https://huggingface.co/datasets/Filo-White/ITAMed Licence The dataset, English translations, annotations, documentation, and original processing materials are released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0). The original Italian examination texts are official acts of the Italian State and are attributed to the Italian Ministry of University and Research. Contacts Filippo Bianchinibianchini@diag.uniroma1.it Edoardo Bianchinie.bianchini@policlinicocampus.it



