遇见数据集

Accuracy and Completeness of AI Chatbot Responses for Leopard Gecko (<i>Eublepharis macularius</i>) Husbandry: An Exploratory Comparison of ChatGPT-4o and Gemini 2.5 Pro

收藏
NIAID Data Ecosystem2026-05-10 收录
官方服务:

资源简介:

Large Language Models (LLMs) like OpenAI’s ChatGPT-4o and Google’s Gemini 2.5 Pro are increasingly used to retrieve care advice on specialized topics including exotic pet husbandry, raising concerns about accuracy and animal welfare. This study, to our knowledge, is the first direct comparison of these models for husbandry knowledge on the widely kept reptile leopard gecko (Eublepharis macularius). Both were queried with 21 questions (biology, care, health) using two prompting styles: questions asked individually in separate chats and all questions asked within a single chat; responses were scored for accuracy and completeness. Both provided generally accurate information (≥2/3) without dangerous errors or “hallucinations.” However, completeness varied; Gemini 2.5 Pro providing more detail. Prompting style was critical: individual prompts yielded far more comprehensive responses (average words: ChatGPT 136, Gemini 387) than combined prompts (ChatGPT 12, Gemini 64), which were sometimes incomplete, especially for ChatGPT-4o. Furthermore, the models demonstrated potential data biases – such as omitting commercial diets while recommending only larger-than-standard enclosures. LLMs can be useful starting points, but critical evaluation, effective prompting, and expert consultation remain essential to ensure animal welfare.

诸如OpenAI的ChatGPT-4o与谷歌的Gemini 2.5 Pro等大语言模型(Large Language Models,LLMs),正愈发广泛地被用于检索涵盖异宠饲养在内的专业主题的护理建议,由此引发了人们对其内容准确性及动物福利的担忧。据我们所知,本研究是首次针对广受饲养的爬行动物豹纹守宫(Eublepharis macularius)的饲养知识,对上述两款模型开展直接对比。研究团队采用两种提示风格,向两款模型各提出21个涵盖生物学、饲养、健康领域的问题:一种是在独立对话中逐一提问,另一种是在单一场对话中一次性提出全部问题;随后对模型回复的准确性与完整性进行评分。两款模型均生成了整体准确率不低于三分之二的准确信息,未出现危险错误或‘幻觉’现象。但二者的回复完整性存在差异:Gemini 2.5 Pro提供的细节更为丰富。提示风格对结果影响显著:逐一提问的方式生成的回复远比一次性合并提问的回复更为详尽(平均字数:ChatGPT-4o为136,Gemini 2.5 Pro为387),后者有时会出现回答不完整的情况,尤其在ChatGPT-4o上表现更为突出。此外,两款模型均存在潜在的数据偏差,例如在推荐使用大于标准尺寸的饲养箱的同时,却未提及商业化饲料。大语言模型可作为有价值的信息起点,但为保障动物福利,对生成内容进行严格评估、采用有效的提示策略以及咨询专业人士,仍是必不可少的环节。

创建时间:
2025-12-24
二维码
社区交流群
二维码
科研交流群
商业服务