FideAI/fmg-bench
收藏资源简介:
FMG-Bench v1 是一个包含120个场景的基准测试数据集,专门用于评估大型语言模型在神学分类和牧灵指导相关背景下的行为。该数据集旨在帮助研究人员和工程师研究模型如何处理涉及教义、传统、道德指导、用户偏好、基础知识和升级边界等面向信仰的问题。数据集包含120个基础场景和37个扰动变体,文件以JSONL格式提供,包括基准分割和示例样本。每个场景包含用户询问、分类级别、传统范围、教义位置、评分权重、预期行为、禁止失败模式以及可选的扰动变体等信息。该数据集仅用于评估目的,不认证牧灵权威、临床安全性、法律充分性或部署准备情况。
FMG-Bench is a 120-scenario benchmark for evaluating large language model behavior in theological triage and pastoral-guidance-adjacent contexts. This release contains the open v1 benchmark corpus: 120 base scenarios with 37 perturbation variants. It is intended for researchers and engineers studying how models handle faith-facing questions involving doctrine, tradition, moral guidance, user preference, grounding, and escalation boundaries. The dataset includes splits for benchmark and example sample, with each scenario containing fields such as user_ask, triage_level, tradition_scope, doctrine_loci, scoring weights, expected_behaviors, disallowed_failure_modes, and optional perturbation variants. It is an evaluation dataset and should not be interpreted as endorsement for real-world pastoral, counseling, legal, or crisis-support use.



