IVF LLM-Code Benchmark: Replication Package for Evaluating LLM-Generated Code in Clinical Decision Support
收藏资源简介:
Replication package for the book chapter "Evaluating Functional Correctness of LLM-Generated Code for AI-Based Clinical Decision Support: An IVF Case Study with Human-in-the-Loop Validation" (S. Sapakova, A. Sapakov), submitted to the edited volume "AI in Software Engineering: Ensuring Quality, Reliability, and Security" (Springer, Studies in Computational Intelligence). The package evaluates code generated by four large language models (Claude Opus 5.5, GPT-6 Sol, Gemini 3.1 Pro, Qwen3-Coder) for 30 tasks derived from a real IVF decision support pipeline: data preprocessing, feature engineering, outcome definition, data splitting, model training, SHAP explanation, and protocol recommendation. Contents: task specifications and the exact prompts; a synthetic sample (300 rows) reproducing the structure of the clinical registry; reference implementations; unit tests built from real-data formats; all 600 generated solutions with response metadata; the notebook used to query the models via OpenRouter; statistical analysis scripts; result tables (test outcomes with error classes E1–E8, pass@k per task, per-solution real-data impact, 72 disputed raw values); figures. The package contains no patient data. The clinical registry (International Clinical Center for Reproductology "PERSONA", Almaty, Kazakhstan) cannot be shared; its research use was approved by the Ethics Committee of the International Information Technology University (IITU), Minutes No. 3, 8 September 2026.



