sentence and clause length distribution
收藏资源简介:
The present study aims to employ the probabilistic distribution of sentence and clause lengths to distinguish translation directionality. Specifically, it addresses three research questions: (1) whether sentence length probability distributions can better discriminate translation directions than average sentence length, (2) whether clause length probability distributions can better discriminate translation directions than average clause length, and (3) which distributional pattern, rank-frequency (R-F) or length-frequency (L-F), is more sensitive to translation direction. Analysis reveals that mean sentence and clause lengths are unreliable metrics for distinguishing translation directions. In contrast, the parameters of the sentence length L-F distribution, fitted as the Extended Positive Negative Binomial (EPNB) model (<i>k</i>, <i>p</i>), and those of the clause length R-F distribution, fitted as the Hyperpoisson model (<i>a</i>, <i>b</i>), exhibit strong discriminative power for translation directionality. The findings demonstrate that probabilistic distributional patterns can better capture and characterize the nuanced linguistic features of translated texts than simple mean values. The study highlights the effectiveness of the data-driven, probabilistic and quantitative linguistic approach in analyzing the sophisticated translation phenomena.
本研究旨在借助句子与分句长度的概率分布特征,判别翻译方向性。具体而言,本研究聚焦三个研究问题:(1) 相较于平均句长,句子长度概率分布是否能更有效地判别翻译方向;(2) 相较于平均分句长,分句长度概率分布是否能更有效地判别翻译方向;(3) 哪种分布模式——词次-频率(rank-frequency,R-F)或长度-频率(length-frequency,L-F)——对翻译方向性更为敏感。分析结果表明,平均句长与平均分句长并非区分翻译方向的可靠指标。与之相对,拟合为扩展正负二项分布(Extended Positive Negative Binomial,EPNB)模型(参数为k、p)的句子长度L-F分布参数,以及拟合为超泊松模型(Hyperpoisson)(参数为a、b)的分句长度R-F分布参数,均展现出极强的翻译方向判别能力。研究结果显示,相较于简单的均值指标,概率分布模式能够更精准地捕捉并刻画翻译文本中细微的语言特征。本研究凸显了以数据为驱动、基于概率与量化的语言学分析方法,在剖析复杂翻译现象时的有效性。




