MammoTab 25: A Large-Scale Dataset for Semantic Table Interpretation – Training, Testing, and Detecting Weaknesses
收藏资源简介:
MammoTab 25 is a large‑scale, richly‑annotated benchmark designed to advance research on Semantic Table Interpretation (STI) and to evaluate the reasoning abilities of modern Large Language Models (LLMs). Scale and origin – The corpus contains 838930 tables automatically extracted from 63 million English‑language Wikipedia pages. Comprehensive annotations – Every table is accompanied by: Cell-Entity Annotation (CEA), Column-Type Annotation (CTA), Columns-Property Annotation (CPA), Four ready‑to‑use prompt templates for LLM training and stress‑testing, Fine‑grained metadata capturing column roles (Named‑Entity vs Literal), NIL flags, header/caption context, and structural statistics. Challenge coverage – Tagged metadata enables users to isolate and diagnose all key STI challenges, including multi-domain tables, acronyms, aliases, typos, approximate numeric values, and true NIL mentions, making the dataset suitable for both benchmarking and error analysis. Format & access – Tables are stored as CSV files; annotations are provided in separate CSVs following the SemTab format; contextual information is packed in JSON side‑cars. The pipeline for regenerating the dataset is openly available on GitHub at https://github.com/unimib-datAI/mammotab/. The documentation is available at https://unimib-datai.github.io/mammotab-docs/.



