Word segmentation from noisy data with minimal supervision
Word segmentation from noisy data with minimal supervision
批准号:
EP/H050442/1
负责人:
Sharon Goldwater
金额:
$35.9万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --
中文摘要
近年来,自然语言处理领域在机器翻译、文档摘要、主题识别等领域取得了长足的进步。然而,这种成功在很大程度上要归功于在有监督的机器学习方法中使用大量人类注释数据构建的系统。这意味着注释资源较少的语言(低密度语言)没有多少有用的语言技术。因此,自然语言处理研究的一个重要方向是提高我们使用尽可能少的注释数据开发成功系统的能力。对完全无监督系统的研究特别有趣,不仅因为它有可能扩大自然语言处理技术的范围,而且因为它可能揭示人类婴儿在很少或没有明确指导的情况下学习语言的方式。我们建议专注于分词这一特定问题,并开发一种新型的概率模型-无限噪声信道模型,以解决在几乎没有注释数据的情况下的这一问题。分词是指识别文本或语音中的词边界的问题。它出现在许多亚洲语言的NLP系统中,其中的单词不用空格分隔,也适用于学习语言的婴儿,因为大多数口语单词都不用停顿分隔。以前关于无监督分词的工作假设每次出现特定的单词时,它都是以完全相同的方式实现的。然而,对于学习语言的婴儿来说,情况并非如此(因为单词在发音时受到语音变化和噪音的影响),在NLP中也不总是如此(如果输入的文本包含错误,例如由光学字符识别系统产生的错误)。我们的新模型将通过同时执行分词以及噪声和可变性的校正来解决这一缺点,以从未分割的噪声输入中恢复去噪的单词序列。我们计划开发我们的模型的两个不同版本。其中一个将被设计用来纠正语音差异,并将作为人类语言习得的认知模型进行评估。通过这个模型,我们希望深入了解允许婴儿成功地从噪声输入中提取单词的计算机制,特别是证明我们模型中使用的贝叶斯推理技术是对婴儿学习行为的合理解释。我们的模型的第二个版本将被设计为纠正光学字符识别产生的错误,并将作为几种不同语言的分词和纠错NLP应用程序进行评估。我们希望表明,该模型减少了文档中的字符错误数量,同时也产生了成功的切分。我们预计这些改进在低密度语言环境中会特别明显。
英文摘要
In recent years, the field of natural language processing (NLP) has made great advances in a wide range of areas, such as machine translation, document summarization, and topic identification. However, much of this success is due to systems that are built using large quantities of human-annotated data in a supervised machine learning approach. This means that languages with fewer annotated resources (low-density languages) are left without much useful language technology. An important direction in NLP research is therefore to improve our ability to develop successful systems using as little annotated data as possible. Research on completely unsupervised systems is particularly interesting not only for its potential to broaden the reach of NLP technology, but also because it may shed light on the ways in which human infants manage to learn language with little or no explicit instruction.We propose to focus on the particular problem of word segmentation, and to develop a new type of probabilistic model, the infinite noisy channel model, for solving this problem in settings where little or no annotated data is available. Word segmentation refers to the problem of identifying word boundaries in either text or speech. It arises in NLP systems for many Asian languages, where words are not separated by whitespace, and also for infants learning language, because most spoken words are not separated by pauses. Previous work on unsupervised word segmentation has assumed that every time a particular word occurs, it is realized in exactly the same way. However, this is not the case for infants learning language (since words are subject to phonetic variability and noise in pronunciation), nor is it always true in NLP (if the input text contains errors, such as those produced by an optical character recognition system). Our new model will address this shortcoming by simultaneously performing word segmentation and correction of noise and variability, to recover a sequence of de-noised words from the unsegmented noisy input. We plan to develop two different versions of our model. One of these will be designed to correct for phonetic variability, and will be evaluated as a cognitive model of human language acquisition. With this model, we hope to gain insight into the computational mechanisms that allow infants to successfully extract words from noisy input, and in particular to show that the Bayesian inference techniques used in our model are a plausible explanation of infants' learning behavior. The second version of our model will be designed to correct for errors resulting from optical character recognition, and will be evaluated as a word segmentation and error-correcting NLP application in several different languages. We hope to show that the model reduces the number of character errors in the document while also producing successful segmentations. We expect these improvements to be particularly pronounced in low-density language situations.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21437/interspeech.2015-239
发表时间:
2015
期刊:
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
10.1109/taslp.2016.2517567
发表时间:
2016-03
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[Micha Elsner (Author)]
通讯作者:
Micha Elsner (Author)
DOI:
10.1016/j.csl.2017.04.008
发表时间:
2016-06
期刊:
ArXiv
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
10.3115/v1/p14-1101
发表时间:
2014-06
期刊:
影响因子:
--
作者:
[Stella Frank;Naomi H Feldman;S. Goldwater]
通讯作者:
Stella Frank;Naomi H Feldman;S. Goldwater
共 6 条
Modeling the Development of Phonetic Representations
-
批准号:ES/R006660/1
-
项目类别:Research Grant
-
资助金额:$37.61万
-
财政年份:2018
-
负责人:Sharon Goldwater
-
依托单位:
海外基金