High-speed arithmetic units are foundational to modern computing architectures. However, the performance of these units is fundamentally constrained by the carry propagation delay. In conventional architectures, generating an output bit or digit relies on the temporal propagation of a carry signal from less significant positions. For an operand of length N, traditional Ripple Carry Adders (RCA) exhibit a linear time complexity of O(N), while advanced parallel- prefix structures, such as Carry-Lookahead Adders (CLA) or Wallace trees, reduce this delay to a logarithmic bound of O(log N). [12]. This paper challenges the necessity of this tempo- ral dependency in specific arithmetic domains. We in- troduce a novel Radix-10 scalar multiplication architec- ture—multiplying an N-digit operand by a single-digit multiplier (M ∈ {0, 1, . . . , 9})—that operates in strictly constant time, O(1). Rather than treating the carry as a dynamic signal that must ripple through the circuit, our meth
收藏资源简介:
Abstract—High-speed digital arithmetic is fundamentally bottlenecked by the delay of carry propagation, conventionally bounded by logarithmic time complexity O(log N) in tree-based architectures [10], [12].This paper presents the Dahwi Spatial Harmonic Decomposition Theory (DSHDT), a novel arithmetic framework that reduces the computational complexity of parallel multiplication to a constant time of O(1). By reformu- lating the arithmetic operation as a localized combinatorial mapping rather than a sequential or recursive process, DSHDT ensures that every output digit (yi) is determined exclusively by a spatially bounded neighborhood of inputs. We demonstrate mathematically, through convergence analysis of the residual carry potential, that the carry influence is strictly localized within a maximum radius of ∆ = 2 digits, enabling direct logic evaluation without temporal propagation chains. Hardware synthesis and simulation confirm that the proposed architecture achieves output stabilization within a fixed critical path delay of ≈ 2 · tgate (where tgate is the fundamental logic gate delay), entirely independent of the operand length N. Furthermore, this constant-time performance is achieved while maintaining a strictly linear hardware area scaling of O(N).



