Excess entropy in natural language: Present state and perspectives

Excess entropy in natural language: Present state and perspectives
复制标题

DOI:
10.1063/1.3630929
复制
发表时间:
2011-09-01
期刊:
影响因子:
2.9
通讯作者:
Debowski, Lukasz
Debowski, Lukasz
中科院分区:
数学2区
文献类型:
--
作者:
Debowski, Lukasz

文献摘要

被引文献

相似文献

我们回顾了近年来在理解自然语言中互信息的意义方面的进展。让我们将文本中的单词定义为足够频繁出现的字符串。在以前的几篇论文中,我们已经证明了这样定义的词的幂律分布(a.k.a.如果在长度增加的文本的相邻部分之间存在(算法)互信息的类似幂律增长,则遵守Herdan定律。此外,如果文本按照类似的幂律分布,以高度重复的方式描述一个复杂的无限(算法上)随机对象,则信息的幂律增长也成立。所描述的对象可能是不可变的(如数学或物理常数),也可能随着时间的推移而缓慢演变(如文化遗产)。在这里,我们以一种不那么技术性的方式来反思各自的数学结果。我们还讨论了决定在何种程度上这些结果适用于实际的人类沟通的可行性。(C)2011年美国物理学会。[doi:10.1063/1.3630929]
We review recent progress in understanding the meaning of mutual information in natural language. Let us define words in a text as strings that occur sufficiently often. In a few previous papers, we have shown that a power-law distribution for so defined words (a.k.a. Herdan's law) is obeyed if there is a similar power-law growth of (algorithmic) mutual information between adjacent portions of texts of increasing length. Moreover, the power-law growth of information holds if texts describe a complicated infinite (algorithmically) random object in a highly repetitive way, according to an analogous power-law distribution. The described object may be immutable (like a mathematical or physical constant) or may evolve slowly in time (like cultural heritage). Here, we reflect on the respective mathematical results in a less technical way. We also discuss feasibility of deciding to what extent these results apply to the actual human communication. (C) 2011 American Institute of Physics. [doi: 10.1063/1.3630929]