遇见数据集

mantra_benchmark

收藏
DataONE2026-05-06 更新2026-05-19 收录
官方服务:

资源简介:

MANTRA is a benchmark for evaluating religious and cultural awareness in text-to-image generation models, grounded in Hindu Vedic iconography. The dataset contains 741 prompts spanning 19 Hindu deities across six dimensions of religious sensitivity: food and dietary taboos, clothing and modern attire, intoxication and substances, narrative revisionism, nudity and anatomical testing, and physical appearance and posture. Each (deity, dimension) pair is instantiated as a matched triplet of three prompts that differ by exactly one iconographic element: a Positive arm depicting the canonical norm, a Negative arm in which the canonical element is replaced by a religiously inappropriate violation, and a Neutral arm in which the element is omitted entirely. Sentence structure, setting, and framing are held constant across all three arms by construction, so any difference in model output is attributable solely to the substituted element. This design decouples representational fidelity (does the model know what the deity must look like?) from constraint awareness (does the model resist rendering a religious violation?), enabling independent measurement of each competency. The release consists of two files. prompts.jsonl contains the 741 prompts; each record carries an ikb_ref field that resolves to a specific cell in the Iconographic Knowledge Base. ikb.csv is the Iconographic Knowledge Base itself: 19 deity rows by 12 norm/violation columns (one norm and one violation per dimension). The dataset is released to support evaluation and improvement of cultural-religious sensitivity in generative models.

MANTRA是一款用于评估文本到图像生成模型(text-to-image generation models)宗教与文化认知能力的基准测试集,其构建基于印度教吠陀图像学(Hindu Vedic iconography)。本数据集共包含741条提示词,涵盖19位印度教神祇,覆盖六大宗教敏感性维度:饮食禁忌、服饰与现代着装、成瘾物质与禁忌摄入物、叙事修正主义、裸露与解剖结构测试,以及外形体态与姿势。每一组(神祇,维度)对均被实例化为一组严格匹配的三元提示词组,三者仅在一个图像学元素上存在差异:正向分支(Positive arm)呈现符合正统规范的范式,负向分支(Negative arm)将规范元素替换为宗教意义上不当的违规范式,中性分支(Neutral arm)则完全省略该目标元素。在构建流程中,所有三组提示词的句子结构、场景设定与表述框架均保持一致,因此模型输出的差异仅由替换的元素所导致。该设计将表征保真度(即模型是否知晓该神祇的标准正统形象)与约束认知能力(即模型是否会规避宗教禁忌性呈现)进行解耦,从而能够独立衡量两项核心能力。本次数据集发布包含两个文件:prompts.jsonl 存储全部741条提示词,每条记录均带有ikb_ref字段,可映射至图像学知识库(Iconographic Knowledge Base)的特定单元格;ikb.csv 即为图像学知识库本身,包含19行神祇数据与12列规范/违规项数据(每个维度对应一项规范项与一项违规项)。本数据集的发布旨在支持生成式模型的宗教文化敏感性评估与优化工作。

创建时间:
2026-05-09
二维码
社区交流群
二维码
科研交流群
商业服务