AMuRD
收藏资源简介:
AMuRD是一个专为从收据中提取关键信息和分类设计的多语言标注数据集,由University of Innsbruck和DISCO AI共同创建。该数据集包含47,720个样本,涵盖阿拉伯语和英语,旨在解决零售行业数据分析中的关键挑战。每个样本都包含详细的标注,如商品名称、价格、品牌等,并分类为44个不同的产品类别,以支持更高效的数据分析。AMuRD不仅提供了丰富的数据资源,还通过其详细的标注和分类功能,为研究人员提供了深入理解商品和交易细节的机会,适用于多种应用场景,如自动化业务流程、财务分析和库存管理。
AMuRD is a multilingual annotated dataset specifically designed for key information extraction and classification from receipts, co-created by the University of Innsbruck and DISCO AI. This dataset contains 47,720 samples covering Arabic and English, aiming to address critical challenges in retail industry data analysis. Each sample includes detailed annotations such as product names, prices, brands and others, and is classified into 44 distinct product categories to support more efficient data analysis. Not only does AMuRD provide a rich data resource, but it also offers researchers opportunities to gain in-depth insights into product and transaction details through its detailed annotations and classification functions, and is applicable to multiple application scenarios such as automated business processes, financial analysis and inventory management.



