遇见数据集

nagohachi/union14m_l_cropped

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

union14m_l_cropped是一个WebDataset格式的数据集,基于Union14M-L的训练部分。它包含3.23百万个已裁剪的场景文本识别样本,这些样本来自14个公共STR数据集。数据集仅包含训练分割(easy、normal、medium、hard、challenging),不包括验证分割和评估套件。样本格式为图像和文本对,总大小约20 GB,分为647个分片。数据集的来源包括KAIST、NEOCR、Uber-Text、RCTW等14个数据集,每个数据集有其自己的许可证,使用时需注意商业用途的限制。

union14m_l_cropped is a WebDataset-format mirror of the training portion of Union14M-L, containing 3.23 million already-cropped scene-text-recognition samples drawn from 14 public STR datasets. This repo only includes the training splits (easy, normal, medium, hard, challenging), and excludes the validation split and evaluation suite. The crops are the original axis-aligned-bbox crops from the Union14M authors, repackaged into 647 shards with a total size of ~20 GB. The underlying images are aggregated from 14 distinct source datasets, each with its own license, and the repackaging is released under the MIT License.

提供机构:
nagohachi
二维码
社区交流群
二维码
科研交流群
商业服务