Khwarezm-100: an access-bounded benchmark of medieval Khwarezm-school mathematics
收藏资源简介:
Khwarezm-100 is the benchmark accompanying the paper Access-Bounded Benchmarks (ICOMP 2026). It evaluates whether a language model solves problems from al-Khwarizmi's Kitab al-jabr and companion Khwarezm-school sources by the source's own method. Correctness and methodological fidelity are scored separately; fidelity is rated by blind human annotators on a nine-criterion rubric, with no LLM judge. Version 0.1.0 (pilot) contains the 15 public pilot items with gold answers and gold derivations, the fidelity rubric, the evaluation harness (raw-output collection, blind shuffled annotation sheets, inter-annotator agreement), the script reproducing the paper's power analysis, and a pipeline-check run. No text of copyrighted modern editions is included.



