遇见数据集

mmprolong

收藏
魔搭社区2026-08-23 更新2026-08-23 收录
官方服务:

资源简介:

# MMProLong Server QA Synthesis This is a long-document VQA dataset generated from documents whose per-paper licenses were verified as CC0, CC BY 3.0, or CC BY 4.0. Raw PDFs are **not** included. Each document is rendered once and samples reference the de-duplicated page images under `images/`. For Chelsea-derived records, PDF bytes were transferred from the third-party `Chelsea707/arxiv-cs-2020-2025-pdfs` mirror at the pinned revision recorded in `data/licenses.jsonl`. The mirror is not treated as a blanket license source. Title, authors, arXiv URL, and license were verified paper by paper using official arXiv metadata. Retain the attribution fields when redistributing the page images. The dataset contains `extract_single`, `extract_multi`, and `reasoning` tasks. Answers were generated by an OpenAI-compatible multimodal teacher endpoint and passed local evidence/page/duplicate validation. Teacher model: `gpt-5.6-luna`; prompt version: `mmprolong-quality-v2`. Target task quotas are `extract_single=1200`, `extract_multi=1200` and `reasoning=600`. Source documents are filtered to `32-50` pages. Check `data/licenses.jsonl` before redistributing any document or image assets. Each asset remains governed by its row-level license and attribution requirements; the synthetic QA annotations do not replace or broaden those source licenses. The JSONL paths are repository-relative. Use `scripts/prepare_qwen.py` or a custom loader to turn `data/train.jsonl` into the exact training input format required by your Qwen2.5-VL trainer.

提供机构:
maas
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务