Data and Code for Sensitivity Analysis of Semantic Clustering Algorithms on Static Word Embeddings
收藏资源简介:
This dataset supports the article “A Comparative Empirical Evaluation of Semantic Clustering Algorithms on Static Word Embeddings” submitted to the International Journal of Information Management Data Insights. The dataset includes curated word lists, sensitivity analysis results, and executable code used to assess the robustness of Phase 1 semantic clustering experiments. Pre-trained static embeddings (GloVe and fastText) are not redistributed due to licensing constraints and were loaded via the Gensim API. The materials enable replication of the reported sensitivity analyses and facilitate further comparative research on semantic clustering robustness.
本数据集用于支撑提交至《国际信息管理与数据洞察期刊》的论文《静态词嵌入(Static Word Embeddings)上语义聚类算法(Semantic Clustering Algorithms)的对比实证评估》。 本数据集包含精选词表、敏感性分析结果,以及用于评估第一阶段语义聚类实验鲁棒性的可执行代码。由于许可限制,预训练静态词嵌入(GloVe与fastText)未随数据集一同分发,仅可通过Gensim应用程序编程接口(API)加载。 本数据集相关材料可复现论文中报道的敏感性分析,并为语义聚类鲁棒性领域的后续对比研究提供支持。




