Evaluating and benchmarking top performing Large Language Models for ecological data mining
收藏资源简介:
This Zenodo release provides all resources needed to reproduce our evaluation of LLM‐driven data extraction in the autonomous ecosystem monitoring case study: Data: CSVs of the 500 abstracts, expert-validated ground‐truth annotations for each target field, and final model outputs for the three core and four additional fields. Scripts: Jupyter notebooks for (1) running batched API calls (LLMFramework.ipynb), (2) computing F₁ scores and cost–time metrics (evaluateAllLLMs.ipynb and evaluateOtherFields.ipynb), and (3) generating all figures from the manuscript (visualizeResults.ipynb and plotF1ScoreGrouped.ipynb). Researchers can reproduce our entire evaluation pipeline—run the scripts against your own API keys and dataset—to validate performance, explore alternative models, or adapt the workflow to new domains.



