LabQAR (<b>Lab</b>oratory <b>Q</b>uestion <b>A</b>nswering with <b>R</b>eference Ranges) Data Set
收藏资源简介:
LabQAR (<b>Lab</b>oratory <b>Q</b>uestion <b>A</b>nswering with <b>R</b>eference Ranges), a manually curated dataset containing multiple-choice questions about 550 lab tests with comprehensive reference ranges sourced from trusted medical resources with annotations on reference ranges, specimen types, and other factors impacting interpretation. We also assess the performance of several large language models (LLMs), including LLaMA 3.1, GatorTronGPT, GPT-3.5, GPT-4, and GPT-4o, in predicting reference ranges and classifying results as normal, low, or high.
LabQAR(全称Laboratory Question Answering with Reference Ranges,即带参考范围的实验室问答数据集)是一份经人工整理的数据集,涵盖针对550项实验室检测的多项选择题,配套源自权威医学资源的完整参考范围,并针对参考范围、样本类型及其他影响结果解读的因素进行了标注。本研究同时基于该数据集评估了多款大语言模型(Large Language Models,缩写LLMs)在参考范围预测及将结果划分为正常、偏低或偏高的分类任务中的性能,涉及的模型包括LLaMA 3.1、GatorTronGPT、GPT-3.5、GPT-4及GPT-4o。



