Proportional Analogy in Written Language Data

Proportional Analogy in Written Language Data
复制标题

书面语言数据中的比例类比

DOI:
10.1007/978-3-319-08043-7_10
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Y. Lepage
Y. Lepage
中科院分区:
--
文献类型:
--
作者:
Y. Lepage

文献摘要

参考文献

被引文献

相似文献

本文的目的是总结多年来比例类比在自然语言处理中的应用研究所取得的一些成果。我们回顾了从欧几里得到现代语言学的概念历史研究中得出的一些基于一般公理的数学形式化。所获得的形式化依赖于两个发音的概念:一致性和比例,并在两个构成性的概念:相似性和连续性。这些概念应用于一系列对象,从集合到符号串,通过多集合和向量,以便获得这些类型的对象中的每一个的数学形式化。由于这些形式化,一些结果,在结构化的语言数据中获得的字符(位图),单词或短句在几种语言,如中文或英文。使用这种仅依赖于形式的形式化的重要一点涉及检索或产生的类比的真实性,即,它们在形式和意义两个层面上是否有效。回顾了这方面的评价结果。还提到了如何符号串的形式化可以应用于两个主要任务,这将对应于索绪尔术语中的“语言”和“言语”:结构化语言数据和生成语言数据。所呈现的结果是从相当大量的语言数据中获得的,例如数千个中文字符或数十万个英语或其他语言的句子。
The purpose of this paper is to summarize some of the results obtained over many years of research in proportional analogy applied to natural language processing. We recall some mathematical formalizations obtained based on general axioms drawn from a study of the history of the notion from Euclid to modern linguistics. The obtained formalization relies on two articulative notions: conformity and ratio, and on two constitutive notions: similarity and contiguity. These notions are applied on a series of objects that range from sets to strings of symbols through multi-sets and vectors, so as to obtain a mathematical formalization on each of these types of objects. Thanks to these formalizations, some results are presented that were obtained in structuring language data by the characters (bitmaps), words or short sentences in several languages like Chinese or English. An important point in using such formalizations that rely on form only, concerns the truth of the analogies retrieved or produced, i.e., whether they are valid on both the levels of form and meaning. Results of evaluation on this aspect are recalled. It is also mentioned how the formalization on string of symbols can be applied to two main tasks that would correspond to ‘langage’ and ‘parole’ in Saussurian terms: structuring language data and generating language data. The results presented have been obtained from reasonably large amounts of language data, like several thousands of Chinese characters or hundred thousand sentences in English or other languages.
DOI: 10.1207/s15516709cog0702_3
发表时间: 1983-01-01
期刊: COGNITIVE SCIENCE
影响因子: 2.5
作者:
GENTNER, D
通讯作者: GENTNER, D