This dataset was extracted from a set of metadata files harvested from the DataCite metadata store (http://search.datacite.org/ui) during December 2015. Metadata records for items with a resourceType
The English corpus comes from the fake news dataset publicly released by Kaggle data analysis competition platform. The Russian corpus comes from the Russian Panorama News Network (Panorama), which fo
Kaleidoscope是一个大规模的多语言多模态考试题库,由Cohere For AI Community创建,包含18种语言的20911个选择题。该数据集旨在评估视觉语言模型在多语言和多模态环境下的表现,涵盖了从高资源语言到低资源语言,以及从数学、社会学到医学和驾驶执照等14个不同的学科领域。数据集通过全球范围内的开源科学合作收集,确保了语言和文化的真实性。