imbWBI: Classification of Business Entities on Multilingual Web - The Main Results
收藏资源简介:
In this research, we proposed and developed, an open source business stakeholder classification system, capable of multi-class single-label hard classification of business entities, according to the products they fabricate. The output is single label result, pointing to the particular industry of the stakeholder. + Summary Spreadsheets with the most relevant findings and research sample data. + TF-IDF Evaluation Contains in total 16 configurations, evaluated in 10-fold cross validation schema, where the same 8 models were ran with page sorting (by text size, desc) at input (of content processing pipeline) and 8 without. Beside the traditional TF-IDF (2 experiments), another 6 modified versions were evaluated: without IDF, with DFC 1.1 and 2.0, and with and without HTML Tag Factors (TW). + Results with CSSRM Cosine SSRM is our customized method for semantic similarity computation. Reports in this folder are performed near and at optimum configuration of the system. + System evaluation Reports and summary spreadsheets on experiments performed for system 10-fold cross validation. + Unstable performance Experiments with different (several sites) sample set, where the system achieved up to F1=0.893 effectiveness, while being unstable because high-number of parallel threads. Morphosyntactic resource interpreter and content decomposition pipeline were producing different results at each run. The results are discarded as non reproducible with single run. ------------------------------------------------------------------- Sample set contains: 5 categories, each having 10 companies (web sites). Specific challenges addressed in this research: - multilingual web content - limited availability of domain-specific training data-sets - heterogeneous linguistic resources of variable quality - absence of production ready and publicly available general semantic lexicons, like WordNet Problems that are addressed by this research: - construction of semantic cloud (non-hierarchical lexicon of semantically related terms) from limited amount of web content - adaptation of similarity computation schema, based on Semantic Similarity Retrieval Model - development of efficient and effective Feature Vector Extraction mechanism, used to reduce number of dimensions in Feature Vector to the number of categories (5) - evaluation of wide range of classification algorithms and configuration parameters: kNN, NaiveBayes, Multiclass SVM and Neural Networks. (17 classifier models are evaluated in every experiment) All software tools (application and the libraries), developed during this research, are published under GNU GPL3 licence, thus available for other researchers and professionals. ---- Goran Grubić Faculty of Organizational Sciences, University of Belgrade, Belgrade, Serbia goran.grubic@koplas.co.rs, +381 62 27 27 55
本研究提出并开发了一款开源商业利益相关者分类系统,该系统可依据商业实体所生产的产品,对其进行多类别单标签硬分类。系统输出为单标签结果,指向该利益相关者所属的特定行业。 + 核心研究发现与研究样本数据汇总电子表格 + TF-IDF(词频-逆文档频率)评估:本次评估共包含16种配置,均采用10折交叉验证(10-fold cross validation)范式开展。其中8个模型在内容处理流水线的输入阶段执行了按文本大小降序排列的页面排序,另外8个未执行该操作。除传统TF-IDF的2组实验外,本次还评估了另外6种改进版本:无IDF版本、DFC 1.1与2.0版本,以及是否添加HTML标签因子(TW)的变体。 + 余弦CSSRM结果:SSRM(语义相似度检索模型,Semantic Similarity Retrieval Model)是本研究定制开发的语义相似度计算方法。本文件夹内的报告均基于系统的近最优与最优配置生成。 + 系统评估报告:包含针对系统10折交叉验证实验的评估报告与汇总电子表格。 + 性能不稳定实验:采用多站点样本集开展的实验中,本系统的F1值最高可达0.893,但由于并行线程数过高导致性能不稳定。形态句法资源解释器与内容分解流水线每次运行的输出结果均存在差异,因此该批次结果因单次运行无法复现而被弃用。 ------------------------------------------------------------------- 样本集详情:共包含5个类别,每个类别对应10家企业(网站)。 本研究攻克的具体挑战包括: - 多语言网页内容处理问题 - 领域专属训练数据集稀缺问题 - 质量参差不齐的异构语言资源问题 - 缺乏可投入生产且公开可用的通用语义词典(如WordNet) 本研究解决的核心问题包括: - 从有限的网页内容中构建语义云(语义相关术语的非层级词典) - 基于语义相似度检索模型适配相似度计算范式 - 开发高效的特征向量提取机制,将特征向量的维度缩减至类别数量(5个) - 评估多种分类算法与配置参数:k近邻(kNN)、朴素贝叶斯(NaiveBayes)、多类别支持向量机(Multiclass SVM)与神经网络。每次实验均评估17个分类器模型。 本研究期间开发的所有软件工具(应用程序与依赖库)均采用GNU GPL3许可证发布,可供其他研究人员与专业人士使用。 ---- 戈兰·格鲁比奇(Goran Grubić) 贝尔格莱德大学组织科学学院,塞尔维亚贝尔格莱德 邮箱:goran.grubic@koplas.co.rs,电话:+381 62 27 27 55




