Assessing latent-structure models in natural language processing with artificial datasets
Assessing latent-structure models in natural language processing with artificial datasets
批准号:
RGPIN-2021-03134
负责人:
Venant, Antoine
金额:
$1.1万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
深度学习现在在自然语言处理(NLP)中无处不在。它将各种任务的性能提升到了新的水平,包括(但不限于):语言建模,句法和语义解析,命名实体识别,情感分析和自然语言理解。深度学习的核心是自动学习和从数据中提取对最终任务有用的特征,从而使开发人员无需将所有语言知识手工制作成计算模型。这样做克服了基于规则的方法遇到的一些困难。它特别提供了一个更容易处理的词汇表项目(利用预先训练的词嵌入),和一个更具表达力的模型集,较少承诺(过于)强的假设(例如,基于形式语言理论的方法中的上下文无关假设)。然而,这些强大的优势也有一些缺点。首先,通常需要非常大量的数据来推断好的模型,在NLP中(特别是在语义和语法方面),这通常是由熟练的注释者在一个耗时且昂贵的过程中产生的。其次,对大量数据的推断计算量很大,并且对环境有很大的影响。最后,深度学习模型对人类来说更难解释。将语言知识,特别是结构知识注入深度学习模型,可能会克服这些局限性。要做到这一点,必须确定如何使深度学习模型意识到语言结构,并找到正确的结构偏差。一个(流行的)研究方向是对语言结构的各个方面进行建模,比如带有潜在变量的语法树。从数据中学习这些模型需要特定的推理算法,通常与方差减少技巧相结合。如果推断出的模型没有提高性能,或者学习到的结构与语言学家的预期有很大的不同(往往是这种情况),那么总是有一个问题,这是否表明未能学习“正确”的结构,或者相反,假定的结构是否没有提供与模型所做的一样好的数据解释。我们建议,为了测试这些推理技术捕捉语言学家理论化的那种结构的严格能力,在人工数据集上评估这些算法。控制实际支配数据生成及其分布的结构应该能够测试这些算法学习给定类型的结构的能力,而不考虑与这些结构的显著性有关的问题,以建模真实的语言数据和其他特定于语言的问题。我们将特别强调推断潜在的语法树(组合)语义解析,并设计我们的人工数据集,以展示特定的结构方面的启发,从语义学家的工作。
英文摘要
Deep learning is now ubiquitous in natural language processing (NLP). It has raised performances to new levels in a variety of tasks, including (but not limited to) : language modeling, syntactic and semantic parsing, named entity recognition, sentiment analysis and natural language understanding. At the heart of deep learning lies the idea of automatically learning and extracting features from the data that are useful to the task at end, thereby saving developers the need to handcraft all linguistic knowledge into computational models. Doing so counters some difficulties encountered by rule-based approaches. It notably offers a much easier handling of out-of-vocabulary items (leveraging pre-trained word embeddings), and a more expressive set of models with less commitments to (overly) strong assumptions (such as context-freeness hypotheses in approaches based on formal language theory, for example). These strong advantages come however with some downsides. First, very large amounts of data are generally required to infer good models, which in NLP (especially in semantics and syntax) is often produced by skilled annotators in a time consuming and expensive process. Second, inference on large amounts of data is computationally heavy and has a strong environmental impact. Finally, deep learning models are much more difficult for humans to interpret. Injecting linguistic knowledge, especially *structural* knowledge into deep learning models might overcome these limitations. To do this, one must determine how to make deep learning models aware of linguistic structure, and find the right kind of structural bias. A (popular) direction of research models aspects of linguistic structure like syntax trees with latent variables. Learning these models from data requires specific inference algorithms, often in combination with variance-reducing tricks. If and when the inferred model do not improve performance, or learned structures differ a lot from linguists' expectations (which tend to be the case), there is always a question of whether this indicates a failure to learn the 'correct' structures, or to the contrary whether the putative structures do not provide as good an explanation of the data as what the model is doing. We propose, in order to test the strict ability of these inference techniques to capture the kind of structures theorized by linguists, to evaluate these algorithms on *artificial* datasets. Controlling the structures actually governing the data generation and their distribution should enable testing the ability of these algorithms to learn a given type of structure, independently from concerns pertaining to the salience of these structures for modeling real linguistic data and other dataset-specific issues. We will put a particular emphasis on inferring latent syntactic trees for (compositional) semantic parsing, and design our artificial datasets to exhibit specific structural aspects inspired from the work of semanticists.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Assessing latent-structure models in natural language processing with artificial datasets
-
批准号:RGPIN-2021-03134
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.1万
-
财政年份:2021
-
负责人:Venant, Antoine
-
依托单位:
Assessing latent-structure models in natural language processing with artificial datasets
-
批准号:DGECR-2021-00145
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Venant, Antoine
-
依托单位:
国内基金
海外基金
基于LMP-1第五跨膜结构域为靶点治疗EB病毒诱导鼻咽癌的药物研发
-
批准号:21602216
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2016
-
负责人:王晓辉
-
依托单位: