遇见数据集

GAMUT

收藏
魔搭社区2026-08-09 更新2026-08-09 收录
官方服务:

资源简介:

# GAMUT🌈: Two-Level Meta-Rubrics for Evaluating Open-Ended Generation 📄 **Paper:** https://arxiv.org/abs/2607.19322 **GAMUT** *(Grounded Assessment of Multimodal Factuality)* is a multimodal **everyday deep-research benchmark** and a **two-level meta-rubric framework** for evaluating open-ended, long-form generation. ![GAMUT overview: a jollof-rice question, its structured meta-rubric, and the compiled binary checklist.](https://huggingface.co/datasets/facebook/GAMUT/resolve/main/assets/figure1.png) **Two-level meta-rubric framework** — resolving a tension in **rubric-based evaluation**. - Judging open-ended generation requires **structure**, because a complete answer often does not decompose into independent, individually required facts: - answers may draw from **open-ended pools** of valid alternatives where no single item is strictly required (the *supporting ingredients* in the figure); - or describe **ordered processes** where the sequence itself matters (the *preparation steps*). - But LLM judges are far more reliable and lower-variance scoring **flat, binary checks** — the very structure needed to *describe* a complete answer is what makes it hard to *grade* one consistently. - GAMUT resolves this with a two-level rubric: a **structured meta-rubric** (left) for authoring the grading criteria — what a complete answer requires — which is then **mechanically compiled** into **flat binary checks** (right) that a judge scores one at a time. **Everyday deep-research benchmark.** - **1,813** image-grounded questions across 10 diverse domains, grounded in real wearable imagery. - Complex, real-world questions that require multi-step information gathering and multi-paragraph answer synthesis. - Questions and rubrics are built through a multi-round human + LLM annotation process, each rubric backed by cited web evidence and verified by expert annotators. **Challenging and discriminative.** - Evaluated 14 proprietary and open-weight models; the best (Gemini 3.1 Pro) reaches only **58.7%**, so the benchmark is far from saturated. - Scores spread widely with large gaps between models, and the ranking is **robust to the choice of judge** (Gemini and Claude judges agree on an identical ranking). ## Dataset fields | field | type | description | |--------------|--------|--------------------------------------------------------| | `session_id` | string | Unique id; also the join key to CRAG-MM for the image. | | `question` | string | The question about the image. | | `rubrics` | struct | The two-level meta-rubric (Answer Critical / Valuable / Context) with source snippets. | ## Splits - **`test`** (1,813 examples) — the main multimodal benchmark; each question is grounded in an image. - **`test_text_only`** (1,806 examples) — a text-only variant, for evaluating models without image input. Same three fields. ```python from datasets import load_dataset ds = load_dataset("facebook/GAMUT", split="test") # multimodal txt = load_dataset("facebook/GAMUT", split="test_text_only") # text-only ``` ## Images Images are **not** included in this dataset. GAMUT references the images in [CRAG-MM](https://huggingface.co/datasets/crag-mm-2025/crag-mm-single-turn-public) by `session_id`. The image-download script and the evaluation harness are in the companion GitHub repository: 👉 **[github.com/facebookresearch/GAMUT](https://github.com/facebookresearch/GAMUT)** - `download_images.py` reconstructs each example's image from CRAG-MM. - `evaluation/` contains the rubric-graded LLM-as-judge evaluation scripts. ## License The GAMUT Dataset is licensed CC-by-NC and is intended for benchmarking purposes only. Third party content pulled from other locations are subject to its own licenses and you may have other legal obligations or restrictions that govern your use of that content. ## Citation <!-- TODO: populate with the final BibTeX entry. --> ```bibtex @misc{chen2026gamut, title={Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: {GAMUT}, a Benchmark for Factual Completeness}, author={Xilun Chen and Zhaleh Feizollahi and Ross Goodwin and Seungwhan Moon and Scott Yih and Pinar Donmez and Babak Damavandi and Luna Dong}, year={2026}, eprint={2607.19322}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2607.19322}, } ```

提供机构:
maas
创建时间:
2026-07-23
二维码
社区交流群
二维码
科研交流群
商业服务