遇见数据集

Accuracy and Completeness of AI Chatbot Responses for Leopard Gecko (<i>Eublepharis macularius</i>) Husbandry: An Exploratory Comparison of ChatGPT-4o and Gemini 2.5 Pro

收藏
Taylor & Francis Group2025-12-24 更新2026-04-16 收录
官方服务:

资源简介:

Large Language Models (LLMs) like OpenAI’s ChatGPT-4o and Google’s Gemini 2.5 Pro are increasingly used to retrieve care advice on specialized topics including exotic pet husbandry, raising concerns about accuracy and animal welfare. This study, to our knowledge, is the first direct comparison of these models for husbandry knowledge on the widely kept reptile leopard gecko (<i>Eublepharis macularius</i>). Both were queried with 21 questions (biology, care, health) using two prompting styles: questions asked individually in separate chats and all questions asked within a single chat; responses were scored for accuracy and completeness. Both provided generally accurate information (≥2/3) without dangerous errors or “hallucinations.” However, completeness varied; Gemini 2.5 Pro providing more detail. Prompting style was critical: individual prompts yielded far more comprehensive responses (average words: ChatGPT 136, Gemini 387) than combined prompts (ChatGPT 12, Gemini 64), which were sometimes incomplete, especially for ChatGPT-4o. Furthermore, the models demonstrated potential data biases – such as omitting commercial diets while recommending only larger-than-standard enclosures. LLMs can be useful starting points, but critical evaluation, effective prompting, and expert consultation remain essential to ensure animal welfare.

以OpenAI的ChatGPT-4o、Google的Gemini 2.5 Pro为代表的大语言模型(Large Language Models,LLMs),如今正被越来越多地用于检索异宠饲养等专业领域的护理建议,这引发了人们对其准确性及动物福利的担忧。据我们所知,本研究首次针对广受欢迎的爬行动物——豹纹守宫(*Eublepharis macularius*)的饲养知识,对上述两款模型进行了直接对比。研究采用两种提示词范式,向两款模型各提出21个涵盖生物学、饲养护理、健康领域的问题:一种是将问题拆分至独立对话会话中单独提问,另一种是将所有问题整合至单一会话中提问;随后对模型回复的准确性与完整性进行评分。两款模型均提供了整体准确率不低于2/3的准确信息,未出现危险性错误或‘模型幻觉’(hallucinations)。但二者的回复完整性存在显著差异:Gemini 2.5 Pro提供的细节更为丰富。提示词范式对结果影响极大:拆分式提问生成的回复更为全面详尽,平均字数分别为ChatGPT 136词、Gemini 2.5 Pro 387词;而整合式提问的回复平均字数仅为ChatGPT 12词、Gemini 64词,且后者时常出现信息缺失,尤以ChatGPT-4o为甚。此外,两款模型还表现出潜在的数据偏见:例如会省略商业饲料的推荐,仅建议使用大于标准尺寸的饲养箱。大语言模型可作为信息检索的有效起点,但为保障动物福利,对其结果进行严谨评估、采用合理的提示词策略,并咨询专业人士仍是不可或缺的关键环节。

提供机构:
Digirolamo, Richard
创建时间:
2025-12-24
二维码
社区交流群
二维码
科研交流群
商业服务