The Probabilistic Representation of Linguistic Knowledge
The Probabilistic Representation of Linguistic Knowledge
批准号:
ES/J022969/1
负责人:
Shalom Lappin
金额:
$58.46万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
In the past twenty-five years work in natural language technology has made impressive progress across a wide range of tasks, which include, among others, information retrieval and extraction, text interpretation and summarization, speech recognition, morphological analysis, syntactic parsing, word sense identification, and machine translation. Much of this progress has been due to the successful application of powerful techniques for probabilistic modeling and statistical analysis to large corpora of linguistic data. These methods have given rise to a set of engineering tools that are rapidly shaping the digital environment in which we access and process most of the information that we use. In recent work (Lappin and Shieber (2007), Clark and Lappin (2011a), Clark and Lappin (2011b)) my co-authors and I have argued that the machine learning methods that are driving the expansion of natural language technology are also directly relevant to understanding central features of human language acquisition. When these methods are used to construct carefully specified formal models and implementations of the grammar induction task, they yield striking insights into the limits and possibility of human learning on the basis of the primary linguistic data to which children are exposed. These models indicate that language learning can be achieved without the sorts of strong innate learning biases that have been posited by traditional theories of universal grammar. Weak biases, some derivable from non-linguistic cognitive domains, and domain general learning procedures are sufficient to support efficient data driven learning of plausible systems of grammatical representation.In the current research I am focussing on the problem of how to specify the class of representations that encode human knowledge of the syntax of natural languages. I am pursuing the hypothesis that a representation in this class is best expressed as an enriched statistical language model that assigns probability values to the sentences of a language. A central part of the enrichment of the model consists of a procedure for determining the acceptability (grammaticality) of a sentence as a graded value, relative to the properties of that sentence and the language of which it is a part. This procedure avoids the simple reduction of the grammaticality of a string to its estimated probability of occurrence, while still characterizing grammaticality in probabilistic terms. An enriched model of this kind will provide a straightforward explanation for the fact that individual native speakers generally judge the well formedness of sentences along a continuum, rather than through the imposition of a sharp boundary between acceptable and unacceptable sentences. The pervasiveness of gradedness in the linguistic knowledge of individual speakers poses a serious problem for classical theories of syntax, which partition strings of words into the grammatical sentences of a language and ill formed strings of words. This research holds out the prospect of important impact in two areas. First, it can shed light on the relationship between the representation and acquisition of linguistic knowledge on one hand, and learning and the encoding of knowledge in other cognitive domains. This work can, in turn, help to clarify the respective roles of biologically conditioned learning biases and data driven learning in human cognition. Second, this work can contribute to the development of more effective language technology by providing insight, from a computational perspective, into the way in which humans represent the syntactic properties of sentences in their language. To the extent that natural language processing systems take account of this class of representations they will provide more efficient tools for parsing and interpreting text and speech.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Towards a Statistical Model of Grammaticality, Proceedings of the 35th Annual Conference of the Cognitive Science Society, Berlin, July-August 2013, pp. 2064-2069.
《走向语法统计模型》,认知科学学会第 35 届年会论文集,柏林,2013 年 7 月至 8 月,第 2064-2069 页。
DOI:
--
发表时间:
2013
期刊:
影响因子:
--
作者:
[Alexander Clark, Gianluca Giorgolo,, Shalom Lappin]
通讯作者:
Shalom Lappin
Jey Han Lau, Alexander Clark, and Shalom Lappin, Predicting Acceptability Judgements with Unsupervised Language Models, Proceedings of the Israeli Seminar in Computational Linguistics, Open University of Israel, June 2015.
Jey Han Lau、Alexander Clark 和 Shalom Lappin,用无监督语言模型预测可接受性判断,以色列计算语言学研讨会论文集,以色列开放大学,2015 年 6 月。
DOI:
--
发表时间:
2015
期刊:
影响因子:
--
作者:
[Jey Han Lau]
通讯作者:
Jey Han Lau
Machine Reading Tea Leaves: Automatically Evaluating Topic Coherence and Topic Model Quality, Proceedings of EACL 2014, Gothenburg, pp. 530-539.
机器阅读茶叶:自动评估主题连贯性和主题模型质量,EACL 2014 年论文集,哥德堡,第 530-539 页。
DOI:
--
发表时间:
2014
期刊:
影响因子:
--
作者:
[Jey Han Lau, David Newman,, Timothy Baldwin]
通讯作者:
Timothy Baldwin
Statistical Representation of Grammaticality Judgements: The Limits of N-Gram Models, Proceedings of the ACL Workshop on Cognitive Modelling and Computational Linguistics, Sophia, August 2013, pp. 28-36.
语法判断的统计表示:N-Gram 模型的局限性,ACL 认知建模和计算语言学研讨会论文集,Sophia,2013 年 8 月,第 28-36 页。
DOI:
--
发表时间:
2013
期刊:
影响因子:
--
作者:
[Alexander Clark, Gianluca Giorgolo,, Shalom Lappin]
通讯作者:
Shalom Lappin
Jey Han Lau, Alexander Clark, and Shalom Lappin, Grammaticality, Acceptability, and Probability: A Probabilistic View of Linguistic Knowledge
Jey Han Lau、Alexander Clark 和 Shalom Lappin,语法性、可接受性和概率:语言知识的概率观
DOI:
--
发表时间:
2016
期刊:
Cognitive Science
影响因子:
2.5
作者:
[Lau, J.H.]
通讯作者:
Lau, J.H.
共 9 条
海外基金