topographrag-bench
收藏资源简介:
TopoGraphRAG-Bench数据包是一个用于基准测试的严格标注数据集,包含基于201个MMDocIR文档构建的基准注释和源文档工件。核心文件`annotations/benchmark.json`提供了2,024个带有拓扑和证据标注的基准问题。数据集还提供完整的文档源材料,包括原始PDF文件、MMDocIR格式的布局级和页面级JSONL结构化内容、页面图像,以及裁剪后的图表/表格和文本布局图像,以支持多模态检索与生成任务。`manifest.json`文件提供了基准文档ID到文件名的映射。数据以压缩存档形式分发,解压后即可获得完整布局。该数据集适用于评估涉及复杂文档理解、多模态信息检索和基于拓扑结构的问答系统。
The TopoGraphRAG-Bench dataset is a rigorously annotated benchmark dataset designed for performance evaluation. It contains benchmark annotations and source document artifacts constructed from 201 MMDocIR documents. Its core file `annotations/benchmark.json` includes 2,024 benchmark questions annotated with topology and evidence information. The dataset also provides complete source document materials, including original PDF files, layout-level and page-level JSONL structured content in MMDocIR format, page images, as well as cropped chart/table and text layout images to support multimodal retrieval and generation tasks. The `manifest.json` file provides the mapping from benchmark document IDs to their corresponding filenames. The dataset is distributed as a compressed archive, and the full dataset can be accessed after extraction. This dataset is suitable for evaluating systems involving complex document understanding, multimodal information retrieval, and topology-based question answering.




