Multi-Document Reasoning and Decision Support with ChatGPT, Gemini and Claude: A Greek-Language Pilot Study with Cross-Model Evaluation
收藏资源简介:
This repository contains the materials and evaluation reports for an exploratory comparison of ChatGPT, Gemini and Claude on a Greek-language administrative decision-support task. The task concerns a synthetic educational project documented through four interconnected files: a project description, a budget, an implementation schedule and meeting minutes containing approved changes, pending proposals and unresolved issues. Using 7 September 2026 as the scenario’s reference date, each model was asked to produce a concise administrative memo advising whether implementation could proceed as planned and identifying the actions required before launch. Responses were required to rely on the supplied documents, provide specific references where needed and explicitly preserve uncertainty where information was unavailable. The task examines reasoning across multiple documents, including the distinction between approved decisions and unapproved proposals, temporal and numerical consistency, resource adequacy, financial synthesis and the practical usefulness of recommended actions. The deposited materials include:• Four scenario documents, version V2.• The common task prompt.• Three anonymized responses, labelled Α, Β and Γ.• An evaluation prompt containing a predefined ten-point assessment key and scoring instructions.• Separate evaluation reports generated by ChatGPT, Gemini and Claude for the three anonymized responses. The evaluation framework separates the detection of predefined issues from overall response quality. Detection is scored using full, partial or zero credit. Quality is assessed independently across five dimensions: documentary grounding, explanation, proposed actions, distinction between established facts and uncertainty, and clarity and usefulness. The tenth detection criterion assesses the synthesis of financial findings rather than an additional independent inconsistency. The collection supports inspection of both task performance and differences in model-based evaluation. All entities, persons, prices and operational conditions in the scenario are synthetic. The task materials and responses are in Greek. This is a single-scenario pilot study. Its scores describe the deposited responses under the documented conditions and should not be interpreted as general rankings of the models or as evidence of statistically established performance differences. The model-generated evaluation reports are assessment outputs, not independently validated ground truth, and require human review. The materials are shared to support methodological scrutiny, reuse and the development of broader studies of LLM reliability in administrative decision support.



