Supplementary Materials for “Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study”
收藏资源简介:
This supplementary dataset supports the study “Effect of Fitzpatrick Skin Type Prompting on Diagnostic Accuracy in Multimodal Large Language Models: A Within-Image Experimental Study.” It includes the original study protocol, final prompt templates, predefined diagnostic scoring ontology, processed model outputs used for scoring, diagnostic scores, the fully executed analysis notebook, post hoc sensitivity-analysis results, descriptive top-1 and top-3 diagnostic accuracy results, binary malignancy-detection results, and top-1 confusion-matrix analyses with the associated prediction crosswalk and mapping audit. The study evaluated 656 biopsy-confirmed photographs from the Diverse Dermatology Images dataset under four prompt submissions per browser-based model configuration: no-FST, DDI-concordant FST, and two DDI-discordant FST prompts. ChatGPT 5.2 Edu and Gemini 3.1 Pro were evaluated, producing 5,248 model-output evaluations. Histopathologic diagnosis served as the diagnostic reference standard. Fitzpatrick skin type was treated as DDI-assigned metadata rather than independently verified ground truth. The deposited Model Outputs and Scores workbook contains the processed ranked differentials used for scoring; raw pre-trimming responses were not retained. These materials document the study workflow and support reproduction of the reported primary, exploratory, and post hoc analyses.




