Chat Logs Dataset for evaluating LLMs (GigaChat, Llama-17b-maverik, Llama-70b, Qwen-32b) as autonomous experts in the analytic hierarchy process
收藏资源简介:
This dataset contains additional experimental research materials. It includes raw chat transcripts (logs) of interactions with Large Language Models (LLMs) used to evaluate their capability to perform as autonomous experts in the Analytic Hierarchy Process (AHP). The models worked with data from World Hapiness Report 2024 dataset. The repository consists of the following test logs: Test 1: the models worked with data from the top 5 countries in terms of the evaluation of their residents. Test 2: The models worked with data from the top 5 worst countries in terms of their residents' assessments. Test 3: the models worked with contrasting data on the GDP parameter, and the 2 best and 3 worst countries were selected. Test 4: the models worked with contrasting data on the social support parameter, and the 2 best and 3 worst countries were selected. Test 5.1: the models worked with combination of data from tests 1 and 2. Test 5.2: the models worked with combination of data from tests 3 and 4. Language of interactions: English. Language of comments: Russian.



