Word segmentation from noisy data with minimal supervision
Word segmentation from noisy data with minimal supervision
批准号:
EP/H050442/1
负责人:
Sharon Goldwater
金额:
$35.9万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --
中文摘要
近年来,自然语言处理(NLP)在机器翻译、文档摘要和主题识别等广泛领域取得了长足的进步。然而,这种成功在很大程度上要归功于在有监督的机器学习方法中使用大量人类注释数据构建的系统。这意味着带有较少注释资源的语言(低密度语言)没有太多有用的语言技术。因此,NLP研究的一个重要方向是提高我们使用尽可能少的注释数据开发成功系统的能力。对完全无监督系统的研究特别有趣,不仅因为它有可能扩大NLP技术的范围,还因为它可能揭示人类婴儿在很少或没有明确指导的情况下学习语言的方式。我们建议将重点放在分词的特殊问题上,并开发一种新的概率模型,即无限噪声通道模型,用于在很少或没有注释数据的情况下解决这一问题。分词是指在文本或语音中识别词边界的问题。它出现在许多亚洲语言的NLP系统中,这些语言中的单词没有空格分隔,也出现在婴儿学习语言时,因为大多数口语单词没有停顿分隔。以前关于无监督分词的工作假设每次出现一个特定的词时,它都以完全相同的方式实现。然而,婴儿学习语言的情况并非如此(因为单词受到语音变化和发音噪音的影响),在NLP中也并非总是如此(如果输入文本包含错误,例如由光学字符识别系统产生的错误)。我们的新模型将通过同时执行分词和校正噪声和可变性来解决这一缺点,从而从未分割的噪声输入中恢复去噪的单词序列。我们计划开发我们模型的两个不同版本。其中一个将被设计用于纠正语音变异,并将被评估为人类语言习得的认知模型。通过这个模型,我们希望深入了解允许婴儿从有噪声的输入中成功提取单词的计算机制,特别是表明我们模型中使用的贝叶斯推理技术是对婴儿学习行为的合理解释。我们的模型的第二个版本将被设计用于纠正光学字符识别导致的错误,并将被评估为几种不同语言的分词和纠错NLP应用程序。我们希望表明该模型减少了文档中字符错误的数量,同时也产生了成功的分割。我们期望这些改进在低密度语言的情况下特别明显。
英文摘要
In recent years, the field of natural language processing (NLP) has made great advances in a wide range of areas, such as machine translation, document summarization, and topic identification. However, much of this success is due to systems that are built using large quantities of human-annotated data in a supervised machine learning approach. This means that languages with fewer annotated resources (low-density languages) are left without much useful language technology. An important direction in NLP research is therefore to improve our ability to develop successful systems using as little annotated data as possible. Research on completely unsupervised systems is particularly interesting not only for its potential to broaden the reach of NLP technology, but also because it may shed light on the ways in which human infants manage to learn language with little or no explicit instruction.We propose to focus on the particular problem of word segmentation, and to develop a new type of probabilistic model, the infinite noisy channel model, for solving this problem in settings where little or no annotated data is available. Word segmentation refers to the problem of identifying word boundaries in either text or speech. It arises in NLP systems for many Asian languages, where words are not separated by whitespace, and also for infants learning language, because most spoken words are not separated by pauses. Previous work on unsupervised word segmentation has assumed that every time a particular word occurs, it is realized in exactly the same way. However, this is not the case for infants learning language (since words are subject to phonetic variability and noise in pronunciation), nor is it always true in NLP (if the input text contains errors, such as those produced by an optical character recognition system). Our new model will address this shortcoming by simultaneously performing word segmentation and correction of noise and variability, to recover a sequence of de-noised words from the unsegmented noisy input. We plan to develop two different versions of our model. One of these will be designed to correct for phonetic variability, and will be evaluated as a cognitive model of human language acquisition. With this model, we hope to gain insight into the computational mechanisms that allow infants to successfully extract words from noisy input, and in particular to show that the Bayesian inference techniques used in our model are a plausible explanation of infants' learning behavior. The second version of our model will be designed to correct for errors resulting from optical character recognition, and will be evaluated as a word segmentation and error-correcting NLP application in several different languages. We hope to show that the model reduces the number of character errors in the document while also producing successful segmentations. We expect these improvements to be particularly pronounced in low-density language situations.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.21437/interspeech.2015-239
发表时间:
2015
期刊:
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
10.1109/taslp.2016.2517567
发表时间:
2016-03
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[Micha Elsner (Author)]
通讯作者:
Micha Elsner (Author)
DOI:
10.1016/j.csl.2017.04.008
发表时间:
2016-06
期刊:
ArXiv
影响因子:
--
作者:
[H. Kamper;A. Jansen;S. Goldwater]
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
DOI:
10.3115/v1/p14-1101
发表时间:
2014-06
期刊:
影响因子:
--
作者:
[Stella Frank;Naomi H Feldman;S. Goldwater]
通讯作者:
Stella Frank;Naomi H Feldman;S. Goldwater
共 6 条
Modeling the Development of Phonetic Representations
-
批准号:ES/R006660/1
-
项目类别:Research Grant
-
资助金额:$37.61万
-
财政年份:2018
-
负责人:Sharon Goldwater
-
依托单位:
海外基金