TID, VATID
收藏资源简介:
目前没有公开的相机捕获发票图像数据集。为了比较不同的文本检测和关键字定位算法,我们收集了两个包含中国不同省份的出租车和增值税发票的数据集,现公开可用。一个称为出租车发票数据集(简称TID),包含104和140类关键字和字符。请注意,出租车发票的关键字在不同省份之间差异很大,我们收集了来自25个不同省份的样本。另一个称为增值税发票数据集(简称VATID),包含24和57类关键字和字符。对于这两个数据集,我们随机选择了50%的图像作为训练集,其余的分配给测试集。
Currently, there is no publicly available dataset of camera-captured invoice images. To compare different text detection and keyword localization algorithms, we have collected two datasets containing taxi and value-added tax (VAT) invoices from various provinces in China, which are now publicly available. One is called the Taxi Invoice Dataset (TID), which includes 104 and 140 categories of keywords and characters. It is important to note that the keywords on taxi invoices vary significantly across different provinces, and we have collected samples from 25 different provinces. The other dataset is called the Value-Added Tax Invoice Dataset (VATID), which includes 24 and 57 categories of keywords and characters. For both datasets, we randomly selected 50% of the images for the training set, with the remainder allocated to the test set.
数据集概述
数据集名称与类型
- Taxi Invoice Dataset (TID): 包含104和140类别的关键词和字符。
- Value Added Tax Invoice Dataset (VATID): 包含24和57类别的关键词和字符。
数据集内容
- TID: 收集自中国25个不同省份的出租车发票图像,关键词和字符类别因省份而异。
- VATID: 收集自中国不同省份的增值税发票图像。
数据集用途
- 用于比较不同的文本检测和关键词定位算法。
数据集划分
- 随机选取50%的图像作为训练集,剩余50%作为测试集。




