遇见数据集

whodatbo1/bulgarian-state-gazette-ocr

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个多模态对话数据集,包含图像和文本对话信息。每条数据包括消息列表(角色和内容)、图像、期刊号、年份、页码和源PDF文件路径。数据集共约30.8万个样本,分为训练集(24.6万)、验证集(3.08万)和测试集(3.08万)。数据集遵循MIT许可证。

This dataset is a multimodal dialogue dataset containing images and text conversation information. Each entry includes a list of messages (role and content), images, issue number, year, page number, and source PDF file path. The dataset contains approximately 308,000 samples, split into training (246,307), validation (30,788), and test (30,789) sets. It is licensed under MIT.

提供机构:
whodatbo1
二维码
社区交流群
二维码
科研交流群
商业服务