LOKI
收藏资源简介:
LOKI数据集由中山大学和上海人工智能实验室等机构联合创建,旨在评估大型多模态模型在检测合成数据方面的能力。该数据集包含视频、图像、3D、文本和音频五种模态,共计18,000条问题,覆盖26个详细子类别。数据集的创建过程包括使用多种合成模型生成高质量数据,并通过精细的异常标注进行分级。LOKI数据集主要应用于合成数据检测领域,旨在解决未来互联网中合成数据泛滥带来的真实性鉴别难题。
The LOKI Dataset, co-developed by Sun Yat-sen University, Shanghai AI Laboratory and other institutions, aims to evaluate the capability of large multimodal models in detecting synthetic data. This dataset covers five modalities including video, image, 3D, text and audio, with a total of 18,000 questions spanning 26 detailed subcategories. The development process of the LOKI Dataset entails generating high-quality data via multiple synthetic models, followed by hierarchical grading through fine-grained anomaly annotation. Primarily applied in the field of synthetic data detection, the LOKI Dataset is designed to tackle the challenge of authenticity verification caused by the widespread proliferation of synthetic data in the future Internet.




