Empirical Estimates of Adaptation: The chance of Two Noriegas is closer to p/2 than p2

Empirical Estimates of Adaptation: The chance of Two Noriegas is closer to p/2 than p2
复制标题

适应的经验估计:两个 Noriegas 的概率更接近 p/2 而不是 p2

DOI:
--
复制
发表时间:
2000
期刊:
International Conference on Computational Linguistics
影响因子:
--
通讯作者:
Kenneth Ward Church
Kenneth Ward Church
中科院分区:
--
文献类型:
--
作者:
Kenneth Ward Church

文献摘要

被引文献

相似文献

重复是很常见的。自适应语言模型允许在看到文本的几个单词后改变或适应概率,它被引入语音识别中以解释文本的连贯性。假设一份文件提到了诺列加一次。他再次被提及的可能性有多大?如果第一个实例的概率为p,那么在标准(词袋)独立性假设下,两个实例的概率应该为p2,但我们发现概率实际上更接近p/2。第一次提到一个词显然取决于频率,但令人惊讶的是,第二次并不如此。适应性更依赖于词汇内容而不是频率;对实义词(专有名词、技术术语和信息检索的好关键词)的适应性更强,对虚词、陈词滥调和普通名字的适应性更弱。
Repetition is very common. Adaptive language models, which allow probabilities to change or adapt after seeing just a few words of a text, were introduced in speech recognition to account for text cohesion. Suppose a document mentions Noriega once. What is the chance that he will be mentioned again? If the first instance has probability p, then under standard (bag-of-words) independence assumptions, two instances ought to have probability p2, but we find the probability is actually closer to p/2. The first mention of a word obviously depends on frequency, but surprisingly, the second does not. Adaptation depends more on lexical content than frequency; there is more adaptation for content words (proper nouns, technical terminology and good keywords for information retrieval), and less adaptation for function words, cliches and ordinary first names.