Datasets for the manuscript: Evaluating multimodal commercial and open-source large language models for dynamical astronomy
收藏资源简介:
The datasets used to evaluate the performance of different large language models on astronomical images (mean-motion and secular resonances). There are four datasets: 'RB–TEST — a test dataset containing eight clear examples that can be used for a quick check of the given model; RB–PILOT — a dataset containing 50 images of different behaviors with simplified outputs (resonant or non-resonant) to compare the results with previous studies; RB–SMALL — a pilot dataset containing 50 images of different behaviors with controversial examples; RB–FULL containing 450 images of all possible behaviors.' (Smirnov & Carruba, 2026-7) Relevant theoretical background can be found in the papers provided in Related works section. If you use this dataset, please consider referencing: Smirnov & Carruba (2026). Submitted.



