act13/AVA-Bench
收藏资源简介:
AVA-Bench是一个用于评估视觉基础模型的诊断性基准,通过原子视觉能力(如定位、计数、OCR、空间理解、深度估计、颜色识别、纹理识别和细粒度识别)来测试模型的基本感知技能。该数据集将视觉感知解耦为14种原子视觉能力,每种能力都包含分布匹配的训练和评估分割,使研究人员能够测量视觉基础模型的优势和不足,并构建能力级别的“能力指纹”以进行模型比较和选择。本版本包含AVA-Bench的训练分割,数据以Parquet分片存储,图像字节嵌入文件中,每个示例包括图像、唯一标识符和指令调优风格的对话字段。数据集适用于视觉基础模型和视觉语言系统的研究,包括训练或指令调优、诊断模型缺陷、通过能力级别性能比较模型以及构建跨视觉能力的平衡训练混合。
AVA-Bench is a diagnostic benchmark for evaluating Vision Foundation Models (VFMs) through Atomic Visual Abilities (AVAs): fundamental perceptual skills such as localization, counting, OCR, spatial understanding, depth estimation, color recognition, texture recognition, and fine-grained recognition. The dataset disentangles visual perception into 14 atomic visual capabilities, each with distribution-matched training and evaluation splits, allowing researchers to measure where a VFM excels or fails and to construct capability-level ability fingerprints for model comparison and selection. This release contains the training split of AVA-Bench, stored as Parquet shards with image bytes embedded in the file, and each example includes an image, a unique identifier, and instruction-tuning style conversation fields. It is intended for research on vision foundation models and vision-language systems, suitable for training or instruction-tuning, diagnosing model deficiencies, comparing models through capability-level performance, and constructing balanced training mixtures across visual abilities.



