Replication package for: An Independent Evaluation of AI Scientist v2 on Biomedical Research Tasks
收藏资源简介:
Replication package for an independent, pre-registered evaluation of Sakana AI's AI Scientist v2 on 11 biomedical research prompts (Vallamsetty 2026). This version (v1.1.0) adds the full replication corpus. Version v1.0.0 contained only the source-code snapshot. Archive: paper-0-replication-v1.tar.gz (594 files, 57 PDFs) v2_manuscripts/ — 24 v2-generated manuscript PDFs ideation_records/ — 32 files covering the 20 v2 ideation outputs human_baseline/ — 3 peer-reviewed Journal of Emerging Investigators positive-control comparator papers v1_outputs/ — 3 AI Scientist v1 nanoGPT_lite template outputs scoring_track_A_ai_drafted/ — 125 AI-drafted scoring forms (scorer: hermes-deepseek-pilot-v3.1) scoring_track_B_human_adjudicated/ — 85 human-adjudicated scoring forms bfts_phase3_logs/ — 20 raw BFTS phase3.log files citation_alignment/ — 132 files; 625 citations resolved against Semantic Scholar and OpenAlex numerical_verification/ — 25 pipeline output files scoring_package/ — 118 files: scoring scripts, rubric anchors, form templates preregistration/, collated_data/, manuscript_figures/, MANIFEST_from_repo.md Also included: paper-0-ai-scientist-eval-v1.0.0.zip — the GitHub source snapshot at commit 6dc21bd (code only; superseded by scoring_package/ in the archive above). Evaluated v2 SHA: 96bd51617cfdbb494a9fc283af00fe090edfae48 Pre-registration: osf.io/aqchj, filed prior to running the experimentation phase Source repository: github.com/calnugget/ai-scientist-v2-biomedical-audit Responsible disclosure of two v2 code bugs (F11, F13): SakanaAI/AI-Scientist-v2 issue #133 Both scoring tracks are preserved for every artifact so the AI-drafted first pass and the human adjudication can be compared independently. All quantitative conclusions in the manuscript apply to the pinned configuration and single-reviewer scoring framework evaluated here.



