aigencydev/aigency-v4-evaluation
收藏官方服务:
资源简介:
AIGENCY V4基准评估结果数据集是AIGENCY V4模型卡和白皮书的可验证证据,包含22个基准测试的评估结果。每个基准测试文件夹包含一个`scored.jsonl`(每项预测、黄金答案、分数)和一个`summary.json`(聚合准确性与Wilson 95%置信区间)。数据集支持土耳其语和英语,涵盖文本生成、多项选择、问答和图像文本到文本等任务类别。
The AIGENCY V4 Benchmark Evaluation Results dataset is the verifiable evidence behind the AIGENCY V4 model card and whitepaper, containing evaluation results across 22 benchmarks. Each benchmark folder includes a `scored.jsonl` (per-item predictions, gold answers, scores) and a `summary.json` (aggregate accuracy with Wilson 95% CI). The dataset supports Turkish and English languages and covers task categories such as text-generation, multiple-choice, question-answering, and image-text-to-text.
提供机构:
aigencydev


