hsclassify-micro-dataset
收藏资源简介:
HSClassify微训练数据集支持海关和贸易工作流程中的多语言HS代码分类。该数据集结合了HS命名记录(6位级别和层次结构上下文)、映射到HS代码的合成产品描述以及用于用户界面和潜在空间分析的人类可读章节/类别标签。数据集包含以下文件:训练数据索引CSV、HS表格快照、HS代码参考JSON和来源归属文件。数据字段包括产品描述文本、6位HS代码目标、章节描述文本、章节ID、标准化人类可读类别标签、HS描述和语言代码。核心HS命名内容来源于`datasets/harmonized-system`项目,上游数据许可证为ODC公共领域奉献和许可证(PDDL)v1.0。项目添加的合成文本和标准化标签在本项目的MIT许可证下发布。当前版本的语言平衡偏向英语,合成文本模式可能未覆盖所有商业短语边缘情况,此数据集仅供研究/原型设计使用,不构成法律海关建议。
The HSClassify micro-training dataset enables multilingual HS code classification for customs and trade workflows. This dataset integrates HS nomenclature records (6-digit level and hierarchical context), synthetic product descriptions mapped to HS codes, and human-readable chapter/category labels for user interface and latent space analysis. The dataset contains the following files: training data index CSV, HS table snapshot, HS code reference JSON, and source attribution file. Its data fields include product description text, 6-digit HS code target, chapter description text, chapter ID, standardized human-readable category labels, HS descriptions, and language codes. The core HS nomenclature content is sourced from the `datasets/harmonized-system` project, with the upstream data license being the ODC Public Domain Dedication and License (PDDL) v1.0. The synthetic text and standardized labels added for this project are released under the MIT license of this work. The current version has a linguistic bias toward English, and the synthetic text patterns may not cover all edge cases of commercial phrases. This dataset is for research and prototyping use only and does not constitute legal customs advice.



