遇见数据集

stepan-omelka/synthetic_omr_500k

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个包含图像和文本转录的多模态数据集,主要用于训练和验证任务。它包含447,656个训练样本和49,788个验证样本,总数据大小约为48.0 GB,下载大小约为47.9 GB。每个样本由图像和对应的文本转录组成,但数据集的具体描述、来源、应用领域或目标在README中未提及。

This dataset is a multimodal dataset containing images and text transcriptions, primarily used for training and validation tasks. It includes 447,656 training samples and 49,788 validation samples, with a total dataset size of approximately 48.0 GB and a download size of approximately 47.9 GB. Each sample consists of an image and its corresponding text transcription, but specific details about the datasets description, source, application domain, or objectives are not provided in the README.

提供机构:
stepan-omelka
二维码
社区交流群
二维码
科研交流群
商业服务