ADAM: Benchmarking and Demonstration Data for Language-Agent-Guided Physical Experimentation
收藏资源简介:
This repository contains benchmarking data and source figures supporting "Balancing Autonomy and Oversight in Language-Agent-Guided Physical Experimentation," which introduces ADAM (Autonomous Decision-making Agent for Materials), a multi-agent large language model framework for safe, human-in-the-loop autonomous experimentation on electron and ion beam instrumentation. Contents include: An expert-curated set of 20 question–answer pairs used to benchmark literature grounding, analytical actionability, hardware safety, and procedural efficiency (qa_pairs.jsonl) Emulator environment images used for closed-loop procedural automation evaluation, including approximately 12,652 PFIB images spanning focus and stigmation conditions Raw source data underlying Figures 2–4, including PFIB demonstration images and model outputs, and TEM demonstration images and model outputs.



