corral_lfm_binomial_results
收藏资源简介:
# *Corral* – LFM Binomial IRT Results <div align="center">  [](https://lamalab-org.github.io/corral/) [](https://lamalab-org.github.io/corral/docs/) [](https://github.com/lamalab-org/corral) [](https://opensource.org/licenses/MIT) [](https://arxiv.org/abs/2604.18805) [](https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results) *Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments* </div> --- ## 📋 Dataset Summary This dataset is part of the *Corral* collection accompanying the paper [*AI scientists produce results without reasoning scientifically*](https://arxiv.org/abs/2604.18805). It contains the **fitted parameters of a binomial Item Response Theory (IRT) model** estimated from agent evaluation runs across all **8 Corral environments**. The dataset is released as a **single config** (`default`). Each row corresponds to a unique combination of *model* and *environment*, and reports the values for every component of the binomial IRT model (e.g., discrimination, difficulty, and latent ability parameters) estimated for that combination. The central finding of this IRT analysis is that **the choice of underlying language model is the dominant source of variance in agent performance**, far outweighing the contribution of the agent scaffold (ReAct, ToolCalling, LLMPlanner, Reflection, etc.). This resource is designed for **psychometric and variance-decomposition analyses** of LLM-based scientific agents. ### 🎯 Supported Uses - 📊 Quantifying the relative contributions of model and scaffold to agent performance - 🔁 Reproducing and extending the IRT analyses reported in the paper - 📐 Psychometric evaluation and calibration studies of frontier LLMs on scientific tasks - 🔍 Identifying environment-specific difficulty and model-specific ability parameters --- ## 🧪 About *Corral* [*Corral*](https://lamalab-org.github.io/corral/) is a framework for the *science of agents and agents for science*. It provides a microservice architecture that **decouples agents from environments** via a client–server design (REST API), ensuring flexibility, reproducibility, and robust isolation. - 🌍 **Environments** define the task space, available tools, and observable feedback — from chemistry labs to HPC clusters. - 🤖 **Agents** are modular LLM-based entities supporting scaffolds such as ReAct, ToolCalling, LLMPlanner, and Reflection. - 📝 **Tasks** define problems to solve, complete with scoring functions. Tasks can be chained into TaskGroups for complex multi-stage challenges. *Corral* currently ships **8 environments**, **97 tools**, **115 tasks**, and **786 subtasks** spanning chemistry, physics, and materials science. ### 🌍 Environments | Environment | Description | 🔧 Tools | 📝 Tasks/scope | 🔭 Scopes | ⏱️ Avg. trace length | |---|---|:---:|:---:|:---:|:---:| | 🧫 **Inorganic Qualitative Analysis** | Identify unknown cations in solution through systematic wet-lab procedures (reagent addition, flame tests, pH measurement, centrifugation, etc.). Observations are computed from thermodynamic data. Three scopes progressively increase the number of candidate ions. | 14 | 10 | 3 | 39.4 | | ⚡ **Circuit Inference** | Recover the topology and component values of a hidden resistor network from pairwise resistance measurements. Tools provide series/parallel calculations, delta-wye transforms, and circuit validation. | 9 | 6 | 1 | 15.0 | | 🔭 **Spectroscopic Structure Elucidation** | Determine the molecular structure of an unknown compound by requesting and interpreting spectroscopic data (MS, NMR, HSQC, IR) alongside reference databases for chemical shifts and isotope distributions. | 16 | 20 | 2 | 15.1 | | 🧬 **Retrosynthetic Planning** | Design multi-step synthetic routes to target molecules under cost, step-count, and commercial-availability constraints, using a template catalogue and functional-group detection tools. | 15 | 8 | 3 | 25.5 | | 🤖 **ML-based Property Prediction** | Assemble a complete ML pipeline to predict formation energies of material polymorphs using data from the Materials Project, covering feature engineering, XGBoost training, and cross-validation. | 14 | 3 | 1 | 16.6 | | 🔬 **AFM Experiment Execution** | Analyze and interpret atomic force microscopy data for nanoscale surface characterization, including topographical and mechanical property measurements. | 6 | 1 | 4 | 26.3 | | ⚛️ **Molecular Simulation** | Design and execute molecular dynamics simulations with LAMMPS to predict materials properties, covering the full workflow from crystal structure retrieval to force-field queries and log analysis. | 8 | 2–3 | 2 | 30.4 | | 🏗️ **Adsorption Surface Construction** | Build adsorbate–slab configurations from bulk crystal structures for heterogeneous catalysis studies, integrating Materials Project retrieval, slab generation, and adsorption-site enumeration. | 15 | 3 | 1 | 19.6 | --- ## 🗂️ Dataset Structure ### Configs This dataset is released as a **single config** (`default`). All model–environment parameter estimates are contained within this one configuration. ### Data Splits The config exposes a single `train` split. ### Data Instances Each row corresponds to a unique **model x environment** combination and contains the fitted values for every component of the binomial IRT model, including discrimination, difficulty, and latent ability parameters estimated for that combination. --- ## 🏗️ Dataset Creation ### Curation Rationale This dataset was created as part of *Corral* to enable psychometric analysis of LLM-based scientific agents, specifically to decompose the sources of variance in agent performance using a binomial IRT framework. The study tests whether agent performance is primarily driven by the underlying language model or by the choice of scaffold. ### Source Data Parameters are derived by fitting a binomial IRT model to agent evaluation outcomes on *Corral* benchmark tasks, covering all evaluated language models and all 8 environments. The source evaluation runs span multiple scaffold types (ReAct, ToolCalling, LLMPlanner, Reflection). --- ## 🔗 Relation to Other Corral Artifacts This dataset is one component of the broader *Corral* release and is best interpreted together with the matching task definitions, execution traces, reports, aggregate results, and reasoning annotations available in the [*Corral* collection](https://huggingface.co/collections/jablonkagroup/corral). --- ## 📄 Citation ```bibtex @article{ríos-garcía2026ai, title = {AI scientists produce results without reasoning scientifically}, author = {Martiño Ríos-García and Nawaf Alampara and Chandan Gupta and Indrajeet Mandal and Sajid Mannan and Ali Asghar Aghajani and N. M. Anoop Krishnan and Kevin Maik Jablonka}, year = {2026}, journal = {arXiv preprint arXiv: 2604.18805} } ``` ## 📜 License This dataset is released under the [MIT License](https://opensource.org/licenses/MIT). ## Changelog ### 2026-04-22 - Initial release of the dataset card.



