Dataset of the study: "Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard"
收藏资源简介:
This dataset contains the 30 questions that were posed to the chatbots (i) ChatGPT-3.5; (ii) ChatGPT-4; and (iii) Google Bard, in May 2023 for the study “Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard”. These 30 questions describe mathematics and logic problems that have a unique correct answer. The questions are fully described with plain text only, without the need for any images or special formatting. The questions are divided into two sets of 15 questions each (Set A and Set B). The questions of Set A are 15 “Original” problems that cannot be found online, at least in their exact wording, while Set B contains 15 “Published” problems that one can find online by searching on the internet, usually with their solution. Each question is posed three times to each chatbot. This dataset contains the following: (i) The full set of the 30 questions, A01-A15 and B01-B15; (ii) the correct answer for each one of them; (iii) an explanation of the solution, for the problems where such an explanation is needed, (iv) the 30 (questions) × 3 (chatbots) × 3 (answers) = 270 detailed answers of the chatbots. For the published problems of Set B, we also provide a reference to the source where each problem was taken from.
本数据集收录了2023年5月,用于题为《数学与逻辑问题场景下的聊天机器人测试:ChatGPT-3.5、ChatGPT-4与Google Bard的初步对比与评估》的研究中,向三款聊天机器人(i)ChatGPT-3.5;(ii)ChatGPT-4;(iii)Google Bard提出的30道测试问题。该30道题目均为具备唯一正确答案的数学与逻辑问题,仅以纯文本形式呈现,无需搭配任何图片或特殊格式排版。题目被划分为A、B两组,每组各15道。其中A组为15道“原创”试题,至少目前无法在互联网上找到完全一致表述的题目;B组为15道“公开”试题,可通过互联网搜索获取,且通常附带解题方案。每道题目会向每个聊天机器人重复提问三次。本数据集包含以下内容:(i)全部30道试题(编号为A01-A15与B01-B15);(ii)每道试题的正确答案;(iii)需配套解题说明的试题对应的解法阐释;(iv)30(道试题)×3(款聊天机器人)×3(次作答)=270份聊天机器人生成的详细作答结果。针对B组的公开试题,本数据集还提供了每道试题的来源引用信息。



