Fetch-A-Set (FAS)
收藏资源简介:
Fetch-A-Set (FAS) 是一个专为立法历史文档分析系统设计的大型基准,旨在解决大规模历史文档检索的挑战。该数据集包含从17世纪至今的文档,总计约40万样本,来源于西班牙的立法文档,覆盖三个世纪。FAS数据集的创建过程涉及使用Mask-RCNN模型识别文档区域,并通过sentencebert编码匹配查询与OCR文本。该数据集主要应用于历史文档分析领域,特别是在文本到图像的检索任务中,旨在通过视觉洞察力提升历史文档分析的效率和准确性。
Fetch-A-Set (FAS) is a large-scale benchmark specifically designed for legislative historical document analysis systems, aiming to address the challenges of large-scale historical document retrieval. This dataset contains documents dating from the 17th century to the present, with approximately 400,000 samples in total. It is sourced from Spanish legislative documents and spans three centuries. The development of the FAS dataset involves utilizing the Mask-RCNN model to recognize document regions, and matching queries with OCR text via Sentence-BERT encoding. This dataset is primarily applied in the field of historical document analysis, particularly for text-to-image retrieval tasks, with the goal of improving the efficiency and accuracy of historical document analysis through visual insights.

- 1Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval计算机视觉中心 2 计算机科学系 巴塞罗那自治大学, 加泰罗尼亚 · 2024年



