遇见数据集

doc_source

收藏
魔搭社区2026-08-20 更新2026-08-23 收录
官方服务:

资源简介:

# doc_source Full document-source corpus for Doc-VQA synthesis. - Source index: `doc_corpus_v1/corpus_index.jsonl` - Total indexed documents: 25,950 - Documents within current size condition: 25,838 - Missing local files at upload time: 0 - Total local PDF bytes: 128,042,509,159 - Within-condition PDF bytes: 116,268,509,199 - Document types: 10 The dataset is organized as: ```text doc_corpus_v1/ corpus_index.jsonl <source>/manifest.jsonl <source>/pdf/*.pdf ``` The corpus index is authoritative for `doc_id`, `doc_type`, `source`, `path`, page count, and the `within_1m` flag. This upload is the whole corpus, not the five-sample review subset. ## Type counts | doc_type | count | |---|---:| | academic_paper | 16,563 | | annual_report_cn | 1,500 | | annual_report_glossy_en | 800 | | congressional_report | 600 | | government_form | 2,598 | | hearing_transcript | 400 | | intl_org_report | 1,238 | | medical_paper | 975 | | technical_standard | 1,152 | | textbook | 124 | ## Source counts | source | count | |---|---:| | arxiv:arxiv_20k_20260724 | 1,938 | | arxiv:arxiv_more2_20260724 | 2,000 | | arxiv:arxiv_more_20260723 | 800 | | arxiv:arxiv_recent5mo_20260725 | 11,705 | | arxiv:arxiv_recent_20260720 | 120 | | cninfo | 1,500 | | edgar_ars | 800 | | govinfo_chrg | 400 | | govinfo_crpt | 600 | | irs | 2,598 | | openstax | 124 | | pmc | 975 | | rfc | 1,152 | | worldbank | 1,238 |

提供机构:
maas
创建时间:
2026-08-20
二维码
社区交流群
二维码
科研交流群
商业服务