An empirical symbolic approach to natural language processing

An empirical symbolic approach to natural language processing
复制标题

DOI:
10.1016/0004-3702(95)00116-6
复制
发表时间:
1996-08-01
影响因子:
14.4
通讯作者:
Velardi, P
Velardi, P
中科院分区:
计算机科学2区
文献类型:
--
作者:
Basili, R;Pazienza, MT;Velardi, P

文献摘要

被引文献

相似文献

自然语言处理(NLP)领域的经验方法通常基于语言的概率模型。这些方法最近得到了普及,因为他们提供了一个更好的语言现象的覆盖范围的索赔。虽然这一主张没有得到完全证明,但在这方面,经验主义方法肯定优于理性主义或象征主义方法。然而,经验方法提供了一个概率,而不是概念,所分析的语言现象的解释。概率系统在真实的应用中确实“工作",这是值得称赞的,但在我们看来,它们本质上无法提供对人类交流机制的洞察,因为输出是由简单的单词或单词群表示的,带有附加的概率。最终,人类分析师必须理解这些数据。在过去的几年里,我们探索了在NLP中结合经验主义和理性主义方法的优势的可能性。我们的目标是定义词汇知识获取的方法,既可扩展,语言上的“吸引力”,也就是说,服从于理论上建立的语言分析。在本文中,我们描述和评估的结果,一个大规模的词汇学习系统,ARIOSTO LEX,使用概率和知识为基础的方法相结合的收购选择限制的话在子语言。我们提出了许多实验数据,从不同领域和不同语言的不同语料库中获得,并表明所获得的词汇数据不仅在自然语言处理中有实际应用,但它们确实是有用的子语言的比较分析。重要的是,ARIOSTO LEX揭示了反复出现的语言现象,这些现象对常用NLP技术的大规模适用性产生了问题影响。
Empirical methods in the field of natural language processing (NLP) are usually based on a probabilistic model of language. These methods recently gained popularity because of the claim that they provide a better coverage of language phenomena. Though this claim is not entirely proved, empirical methods certainly outperform in this regard rationalist, or symbolic, methods. However, empirical methods provide a probabilistic, not conceptual, explanation of the analyzed linguistic phenomena. Probabilistic systems do ''work'' in real applications, and this is meritorious, but in our view they are intrinsically unable to provide insight into the mechanisms of human communication, because the output is represented by plain words, or word clusters, with attached probabilities. Eventually, a human analyst must make sense of these data. In the past few years, we explored the possibility of combining the advantages of empirical and rationalist approaches in NLP. Our objective was to define methods for lexical knowledge acquisition that are both scalable and linguistically ''appealing'', that is, amenable to a theoretically founded analysis of language. In this paper we describe and evaluate the results of a large scale lexical learning system, ARIOSTO LEX, that uses a combination of probabilistic and knowledge-based methods for the acquisition of selectional restrictions of words in sublanguages. We present many experimental data obtained from different corpora in different domains and languages, and show that the acquired lexical data not only have practical applications in NLP, but they are indeed useful for a comparative analysis of sublanguages. Importantly, ARIOSTO LEX shed light on recurrent linguistic phenomena that have a problematic impact on the large-scale applicability of commonly used NLP techniques.