Evaluation data and analysis code — zero-marginal-cost RAG advisory system
收藏资源简介:
Ground-truth question set (this version). This version adds a draft expansion of the ground-truth question set. Contents: - 119 validated answer-keys - human-reviewed, corpus-grounded, and used as the evaluation set in the accompanying paper. All reported paper metrics are based solely on these 119.- 975 draft questions (answer_keys_expansion_draft.jsonl) - synthetically generated draft questions with placeholder reference answers. Explicitly marked validated: false, carry empty source_chunk_ids (no corpus grounding), and were NOT used to compute any reported metric. Released only as a scaffold for future curation. Combined listed total: 1,094 questions (119 validated + 975 The 975 draft records are not a graded evaluation set - placnd reused across questions, and require human review plussource-chunk grounding before any evaluation use. Draft topic breakdown (975): rice 180 · soybeans 135 · dicamba 110 · poultry 95 · weeds 95 · irrigation 85 · diagnostics 80 · fertility 70 · weather 45 · harvest 40 · economics 30 · general 10.



