遇见数据集

Class-Distributed Learning for Multinomial Logistic Regression with High Dimensional Features and a Large Number of Classes

收藏
DataCite Commons2024-07-22 更新2024-08-19 收录
官方服务:

资源简介:

Estimating a high-dimensional multinomial logistic regression model with a larger number of categories is of fundamental importance but it presents two challenges. Computationally, it leads to heavy computation cost. Statistically, it suffers unsatisfactory statistical efficiency. Therefore, how to solve this problem in a computationally and statistically efficient way is of great interest. To tackle these challenges, we have developed a new class-distributed learning algorithm with a rank-reducible coefficient structure. The key innovation here is piecing together two important techniques for distributed computing and improved statistical efficiency. The two techniques are, respectively, dimension reduction and a circular-structured working model. Dimension reduction effectively alleviates the curse of dimensionality due to high dimensional features. A circular-structured working model allows the use of a class-distributed algorithm for distributed computing. To support our new methodology, we develop rigorous asymptotic theory and present extensive numerical experiments. Supplementary materials for this article are available online.

构建类别数较多的高维多项logistic回归模型具有重要的基础理论价值,但同时面临两大挑战。从计算层面来看,该任务会带来高昂的计算开销;从统计层面而言,其统计效率难以令人满意。因此,如何在计算与统计层面均实现高效的解决方案,成为学界广泛关注的研究课题。为应对上述挑战,我们提出了一类带有秩可约化系数结构的新型类别分布式学习算法。本研究的核心创新在于融合了两项分别用于分布式计算与提升统计效率的关键技术:降维技术与循环结构工作模型(circular-structured working model)。降维技术可有效缓解高维特征带来的维数灾难问题;循环结构工作模型则支持通过类别分布式算法开展分布式计算。为验证所提新方法的有效性,我们构建了严格的渐近理论框架,并开展了大量数值仿真实验。本文的补充材料可在线获取。

提供机构:
Taylor & Francis
创建时间:
2024-05-31
搜集汇总
数据集介绍
Class-Distributed Learning for Multinomial Logistic Regression with High Dimensional Features and a Large Number of Classes 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务