Ancient-Bench
收藏资源简介:
Ancient-Bench是一个全面评估古代中国文物文本识别的基准数据集,由华南理工大学与华为技术有限公司联合创建。该数据集包含2700张精心标注的图像,覆盖甲骨、青铜、简牍、帛书、印章、碑刻、摩崖、古籍和书法九种媒介,以及甲骨文、金文、篆书、隶书、楷书、草书和行书七种历史字体,时间跨度超过3000年。数据源自14个国家级文化遗产机构,经格式标准化、质量筛选、区域定位,并制定了符号、字符和解析三种标准化标注规范,确保跨媒介一致性。该基准旨在揭示现有模型在异体字、特殊符号处理及幻觉现象上的系统性不足,为跨时代、跨媒介、跨字体的文物文本识别研究提供统一评估框架。
Ancient-Bench is a benchmark dataset for comprehensive evaluation of ancient Chinese cultural relic text recognition, jointly created by South China University of Technology and Huawei Technologies Co., Ltd. This dataset contains 2700 carefully annotated images, covering nine types of carriers including oracle bone relics, bronze wares, bamboo slips, silk manuscripts, seals, stone steles, cliff inscriptions, ancient books and calligraphy works, as well as seven categories of historical Chinese scripts: oracle bone script, bronze script, seal script, clerical script, regular script, cursive script and running script, with a time span of over 3000 years. The dataset is sourced from 14 national-level cultural heritage institutions, and has undergone format standardization, quality screening and region localization, with three standardized annotation specifications (symbol-level, character-level and parsing-level) formulated to ensure cross-media consistency. This benchmark aims to reveal the systematic limitations of existing models in handling variant Chinese characters, special symbols and hallucination phenomena, and provide a unified evaluation framework for cross-era, cross-media and cross-script research on cultural relic text recognition.
Ancient_Bench 数据集详情
数据集概述
Ancient_Bench 是一个由 SCUT-DLVCLab 团队发布的数据集,其核心内容与“古代”主题相关,旨在为相关领域的研究提供基准测试(Benchmark)支持。
数据集用途
该数据集以基准测试(Benchmark)为主要定位,可用于评估和比较模型在特定任务(推测为古代相关文本、图像或多模态任务)上的表现。
数据集来源与维护
- 发布机构:SCUT-DLVCLab(华南理工大学深度学习与视觉计算实验室)
- 代码仓库:https://github.com/SCUT-DLVCLab/Ancient_Bench
内容说明
当前数据集的 README 文件内容较为简洁,仅包含数据集名称 Ancient_Bench,未提供更详细的字段说明、数据规模、任务类型或评估指标等信息。如需进一步了解数据集的具体构成与使用方式,建议直接访问上述 GitHub 仓库获取详细文档。





