遇见数据集

NusaAksara/NusaAksara

收藏
Hugging Face2025-02-14 更新2025-04-19 收录
官方服务:

资源简介:

该数据集包含了多个子数据集,分别用于图像分割、图像转录(OCR)、图像翻译、图像转写、转录语言识别(LID)、转录翻译和转录转写等任务。每个子数据集都有训练集split,提供了不同的特征,如图像ID、图像URL、高度、宽度、语言、分割信息、转录、翻译、转写和语言标签等。

The dataset consists of multiple sub-datasets for tasks such as image segmentation, image transcription (OCR), image translation, image transliteration, transcription language identification (LID), transcription translation, and transcription transliteration. Each sub-dataset has a training set split and provides various features, including image ID, image URL, height, width, language, segmentation information, transcription, translation, transliteration, and language labels.

提供机构:
NusaAksara
二维码
社区交流群
二维码
科研交流群
商业服务