Data and code for: Are Large Language Models Able to Perform or Assist in Crime Scene Investigations?
收藏资源简介:
Scoring and statistics code, exact prompts, task files with problem identifiers, and every model response with its timing, token counts and cost, for the manuscript Are Large Language Models Able to Perform or Assist in Crime Scene Investigations? (Daniel Attinger, 2026), submitted to Forensic Science International. The study measures whether large language models can carry out the multi-step, abstention-aware reasoning that crime-scene reconstruction requires, using the MuSR benchmark (murder mysteries, object placements, team allocation; Sprague et al., ICLR 2024) with an added I don't know option, plus MMLU-Pro physics and engineering, BBEH and GSM8K as controls. Models: three commercial Claude models (Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5, plus Opus 4.5 and Sonnet 4.5 on one split), the hosted open-weight Qwen3.8-27B at 16-bit and 4-bit precision, Qwen3.8-Max and DeepSeek V4 Pro, the same Qwen3.8-27B weights in a 4-bit GGUF file on a rented RTX 3090 and on the author's gaming notebook, and local Gemma 2 9B, Qwen3.5-9B, Qwen3.5-27B, Qwen2.5-Coder 7B and 14B under Ollama. Contents: eval/ the Python scoring and statistics code (score.py, paired.py, sweep.py, cost-table.py, fresh-compare.py), eval/tasks/ the fixed samples with problem identifiers, eval/runs/ 12,225 recorded model responses (one JSON object per response: model, backend, prompt style, temperature, response text, duration, prompt and output token counts, cost, GPU placement), and the study's own records (RESULTS.md, KNOWN-ISSUES.md, MODEL-PROVENANCE.md, README.md); S1-prompts/ the system instruction and the CoT+ strategy texts verbatim and as JSON; S2-per-problem/ per-problem outcomes for every model and problem type with paired disagreement counts (S2-paired-disagreements.md) and the fresh-item comparison (fresh-compare-2026-09-20.md); S3-preregistered-readings.md the readings of the object-placement runs written before the results were known; S4-generator-selection.md the evidence for the choice of generator for the fresh items, with its quoted pilot stories cut; MANIFEST.txt with the SHA-256 of every file and the reason for every exclusion. The code that sent the problems to the models and ran the campaign is not included; the method is described in full in the manuscript. Licence: the code and documentation are released under the MIT License (LICENSE). The task files are derived from MuSR (MIT), MMLU-Pro (MIT), BBEH (Apache-2.0) and GSM8K (MIT) and keep those licences; see README.md. The fifty freshly generated murder-mystery items used for the contamination test are withheld by the author's decision and are available from the author to the editor under confidentiality.



