jungsin3/ko-sample
收藏资源简介:
ko-sample是一个从KORMo-VLM/Korean-MM-dataset2中抽取的样本数据集,包含五个领域:chart(图表)、documents_ocr_stage1(文档OCR阶段1)、documents_ocr_stage2(文档OCR阶段2)、korean(韩语)和tables(表格)。每个领域有40个样本,每个样本对应40个引用图像和40个提取图像,并分别提供了JSON数据文件和图像目录。该数据集主要用于多模态视觉语言任务,涉及图像和相关文本数据的处理与分析。
The ko-sample is a sample dataset extracted from KORMo-VLM/Korean-MM-dataset2. It encompasses five domains: chart, documents_ocr_stage1, documents_ocr_stage2, korean, and tables. Each domain contains 40 samples, with each sample corresponding to 40 reference images and 40 extracted images. Separate JSON data files and image directories are provided for each domain and sample. This dataset is primarily designed for multimodal vision-language tasks, involving the processing and analysis of image and related text data.




