课题基金 / 基金详情

The Role of Neural Models in (Constrained) Natural Language Generation Built on the mathematical foundations laid out by Markov [1], n-gram language

The Role of Neural Models in (Constrained) Natural Language Generation Built on the mathematical foundations laid out by Markov [1], n-gram language
神经模型在(受限)自然语言生成中的作用建立在马尔可夫 [1] n-gram 语言奠定的数学基础之上
批准号:
2438674
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
建立在马尔可夫[1]奠定的数学基础上,n-gram语言模型赋予一个单词一个依赖于其上下文(它前面的n个单词)的概率分布。虽然这些模型导致了语言技术的进步,如语音识别[2]和机器翻译[3,4],但它们的参数化基数随着上下文长度呈指数增长。这限制了它们的适用性。递归神经网络可以对任意长度的上下文进行建模[5],这导致这些模型在许多语言任务中被广泛采用[6]。克服了递归架构的计算效率限制,[7]中提出的神经“Transformer”架构导致了一系列神经语言模型的开发,这些模型在各种自然语言任务中实现了最先进的性能[8,9,10].该项目旨在推进人工智能(AI)基于自然语言的技术,通过研究如何使用神经语言模型将用户目标纳入自然语言生成。这些目标可以通过与系统的交互来表达,就像对话式AI中的情况一样。受神经模型用于取代复杂模型管道的领域的最新发展的启发[11,12,13],该项目将探索并提供新的方法,使这些系统能够:更好的估计,跟踪并根据用户意图调整或自适应其输出,适应以满足新的用户目标目标也可以通过语言工程师想要嵌入到生成任务中的约束来表达。例如,在语言中生成性别变化,后者取决于人类所指对象的社会性别,这对最先进的自动翻译系统来说是一个挑战[14],最近的工作表明,在神经模型中嵌入这种约束是不平凡的[15]。解决这个问题不仅仅是缓解性别偏见,因为可以采用类似的方法来约束神经系统生成性别中立的翻译。该项目旨在采取广泛的方法来推进用户约束语言生成的最新技术水平,调查:适当的数据源及其最佳表示架构变化,以适应新的/更丰富的输入数据表示新的训练方法,包括新的目标和与其他系统结合的训练新的解码过程,这些解码过程考虑了约束自适应技术。A.(1913)。Essai d 'une recherche statistique sur le texte du roman“尤金奥涅金”illustrant la liaison des epreuve en chain(“尤金奥涅金”文本的统计调查示例,说明了链中样本之间的依赖性“)。Izvistia Imperatorskoi Akademii Nauk(Bulletin de l'Academie Imériale des Sciences de St. Pétersbourg),7,153-162. [2]Povey,D.,& Woodland,P. C. (2002,五月)。最小音素误差和I平滑用于改进的判别训练。在2002年IEEE声学、语音和信号处理国际会议上(第1卷,第110页),I-105)。IEEE。[3]Chiang,D. (2005,六月)。一种基于短语的层次统计机器翻译模型。在计算语言学协会(ACL'05)第43届年会的会议记录中(pp. 263-270)。[4]de Gispert,A.,Iglesias,G.,黑木,G.,R.邦加,E.,& Byrne,W.(2010年)。基于加权有限状态转换器和浅n文法的分层短语翻译。Computational linguistics,36(3),505-533. [5]Mikolov,T.,Karafiát,M.,伯基特湖,Cerricky,J.,& Khudanpur,S.(2010年)。基于递归神经网络的语言模型。在国际语音通信协会第十一届年会上,1045-1048)。[6]Jurafsky,D.,&,Martin,J. H.,(n.d)。用神经网络进行序列处理。语音与语言处理导论
英文摘要
Built on the mathematical foundations laid out by Markov [1], n-gram language models endow a word with a probability distribution that depends on its context (the n words preceding it). While these models led to advances in language technologies as diverse as speech recognition [2] and machine translation [3, 4], the cardinality of their parametrisation grows exponentially with the context length. This limits their applicability. Recurrent neural networks can model arbitrary length contexts [5], which led to the widespread adoption of these models in many language tasks [6]. Overcoming the computational efficiency limitations of recurrent architectures, the neural "transformer" architecture proposed in [7] led to development of a range of neural language models that achieve state-of-the-art performance in a variety of natural language tasks [8, 9, 10].This project aims to advance artificial intelligence (AI) based technologies for natural language by investigating how neural language models can be employed to incorporate user goals in natural language generation. Such goals might be expressed through interaction with the system, as is the case in conversational AI. Inspired by recent developments in the field where neural models have been employed to replace complex model pipelines [11, 12, 13], the project will explore and provide novel methods that will allow these systems to: better estimate, track and ground their output in user intent be adapted or self-adapt to satisfy new user goalsGoals might also be expressed through constraints that a language engineer would like to embed in the generation task. For example, generation of gender inflections in languages where the latter depends on the social gender of a human referent is challenging for state-of-the-art automatic translation systems [14] and recent work has shown that embedding such constraints in neural models is non-trivial [15]. Solving this problem extends beyond gender bias mitigation, as similar approaches might be pursued to constrain neural systems to generate gender-neutral translations. This project aims to take a broad approach to advancing state-of-the-art in user-constrained language generation, investigating: appropriate data sources and their optimal representation architectural changes necessary to accommodate new/richer input data representations new training methodologies, including novel objectives and training in conjunction with other systems novel decoding processes that account for constraints adaptation techniquesReferences[1] Markov, A. A. (1913). Essai d'une recherche statistique sur le texte du roman "Eugene Onegin" illustrant la liaison des epreuve en chain ('Example of a statistical investigation of the text of "Eugene Onegin" illustrating the dependence between samples in chain'). Izvistia Imperatorskoi Akademii Nauk (Bulletin de l'Academie Impériale des Sciences de St.-Pétersbourg), 7, 153-162.[2] Povey, D., & Woodland, P. C. (2002, May). Minimum phone error and I-smoothing for improved discriminative training. In 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing (Vol. 1, pp. I-105). IEEE.[3] Chiang, D. (2005, June). A hierarchical phrase-based model for statistical machine translation. In Proceedings of the 43rd annual meeting of the association for computational linguistics (ACL'05) (pp. 263-270).[4] de Gispert, A., Iglesias, G., Blackwood, G., R. Banga, E., & Byrne, W. (2010). Hierarchical phrase-based translation with weighted finite-state transducers and shallow-n grammars. Computational linguistics, 36(3), 505-533.[5] Mikolov, T., Karafiát, M., Burget, L., Cernocky, J., & Khudanpur, S. (2010). Recurrent neural network based language model. In Eleventh annual conference of the international speech communication association (pp. 1045-1048).[6] Jurafsky, D., &, Martin, J. H., (n.d). Sequence Processing with Neural Networks. In Speech and Language Processing: An Introd
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/2022.findings-acl.223
发表时间: 2022
期刊:
影响因子: --
作者: [Tisha Anders;Alexandru Coca;B. Byrne]
通讯作者: Tisha Anders;Alexandru Coca;B. Byrne
DOI: 10.18653/v1/2021.eancs-1.2
发表时间: 2021
期刊: The First Workshop on Evaluations and Assessments of Neural Conversation Systems
影响因子: --
作者: [Alexandru Coca;Bo-Hsiang Tseng;B. Byrne]
通讯作者: Alexandru Coca;Bo-Hsiang Tseng;B. Byrne
国内基金
海外基金
Neural Process模型的多样化高保真技术研究