HIVMedQA
收藏资源简介:
This dataset supports the findings presented in the article HIVMedQA: Benchmarking large language models for HIV medical decision support. It comprises two components: questions.csv : Contains all the questions, the corresponding gold-standard answers, and their sources. all_questions_answers_scores.csv : Includes the responses generated by LLMs, along with evaluation scores. all_questions_answers_scores_unsupervised.csv: Includes the evaluation scores under the unsupervised setting. all_questions_answers_scores_withRAG_withInfoInGuidelines.csv: Includes the responses generated by LLMs with using RAG and their associated evaluation scores. If you use this dataset in your work, please cite: Cardenal-Antolin, Gonzalo, et al. "HIVMedQA: Benchmarking large language models for HIV medical decision support." arXiv preprint arXiv:2507.18143 (2025).



