DRISHTIKON
收藏资源简介:
DRISHTIKON是一个专注于印度文化的多模态多语言基准数据集,旨在评估生成式AI系统的文化理解能力。该数据集涵盖了15种语言,包括所有州和联邦属地,并包含了超过64,000个对齐的文本-图像对。数据集捕捉了丰富的文化主题,包括节日、服饰、美食、艺术形式和历史遗产等。DRISHTIKON填补了包容性AI研究中的一个重要空白,为推进文化意识强的、多模态的语言技术提供了一个强大的测试平台。
DRISHTIKON is a multimodal, multilingual benchmark dataset focused on Indian culture, designed to evaluate the cultural comprehension capabilities of generative AI systems. It covers 15 languages, spans all Indian states and union territories, and contains over 64,000 aligned text-image pairs. The dataset captures a rich range of cultural themes including festivals, traditional attire, cuisine, art forms, historical heritage and more. DRISHTIKON fills a critical gap in inclusive AI research, serving as a robust testbed for advancing culturally-aware, multimodal language technologies.

- 1DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture印度理工学院帕特纳分校 · 2025年



