遇见数据集

REAL-MM-RAG_FinTabTrainSet_rephrased

收藏
魔搭社区2026-05-13 更新2026-08-30 收录
官方服务:

资源简介:

<style> /* H1{color:Blue !important;} */ /* H1{color:DarkOrange !important;} H2{color:DarkOrange !important;} H3{color:DarkOrange !important;} */ /* p{color:Black !important;} */ </style> <!-- # REAL-MM-RAG-Bench We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world retrieval challenges, we introduce multi-level query rephrasing, modifying queries at three distinct levels—from minor wording adjustments to significant structural changes—ensuring models are tested on their true semantic understanding rather than simple keyword matching. ### Source Paper [REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark](https://arxiv.org/abs/2502.12342) --> ## REAL-MM-RAG_FinTabTrainSet_rephrased We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM. We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. This is the rephrased version of the training set, where each query was rephrased to preserve semantic significance while changing the wording and query structure. The non rephrased version can be found in https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet For more information, see the project page: https://navvewas.github.io/REAL-MM-RAG/ ## Source Paper ```bibtex @misc{wasserman2025realmmragrealworldmultimodalretrieval, title={REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark}, author={Navve Wasserman and Roi Pony and Oshri Naparstek and Adi Raz Goldfarb and Eli Schwartz and Udi Barzelay and Leonid Karlinsky}, year={2025}, eprint={2502.12342}, archivePrefix={arXiv}, primaryClass={cs.IR}, url={https://arxiv.org/abs/2502.12342}, } ```

提供机构:
maas
创建时间:
2025-10-04
二维码
社区交流群
二维码
科研交流群
商业服务