NICE 和 STOPS
收藏资源简介:
NICE 和 STOPS 是由乌尔姆大学创建的两个新数据集,用于短文本分类任务。NICE 数据集基于世界知识产权组织的尼斯分类,将商品和服务分为45类,STOPS 数据集则基于亚马逊产品描述和Yelp商业条目。这两个数据集旨在解决税务审计中商品和服务分类的问题,提供了新的特征,如更短的平均长度和更细粒度的分类能力。创建过程中,数据经过了预处理,包括转换为小写、去除标点符号和随机洗牌。这些数据集在实际应用中,特别是在区分商品和服务方面,具有重要价值。
NICE and STOPS are two novel datasets developed by Ulm University for short text classification tasks. The NICE dataset, which is based on the Nice Classification of the World Intellectual Property Organization (WIPO), categorizes goods and services into 45 classes. The STOPS dataset, by contrast, is built upon Amazon product descriptions and Yelp business listings. These two datasets are designed to address the challenge of goods and services classification in tax audits, and offer new features such as shorter average text length and finer-grained classification capabilities. During the creation process, the data underwent standardized preprocessing steps including lowercasing, punctuation removal, and random shuffling. These datasets hold significant practical value, particularly in scenarios requiring the differentiation between goods and services.



