Assessing latent-structure models in natural language processing with artificial datasets
Assessing latent-structure models in natural language processing with artificial datasets
批准号:
RGPIN-2021-03134
负责人:
Venant, Antoine
金额:
$1.1万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
深度学习现在在自然语言处理(NLP)中无处不在。它将各种任务的性能提升到新的水平,包括(但不限于):语言建模、句法和语义解析、命名实体识别、情感分析和自然语言理解。深度学习的核心思想是自动学习并从数据中提取对最终任务有用的特征,从而节省开发人员将所有语言知识手工制作到计算模型中的需要。这样做可以克服基于规则的方法遇到的一些困难。值得注意的是,它更容易处理词汇表外的项目(利用预训练的词嵌入),并且提供了一组更具表现力的模型,对(过度)强假设的承诺更少(例如,基于形式语言理论的方法中的上下文无关假设)。然而,这些强大的优势也带来了一些缺点。首先,通常需要非常大量的数据来推断良好的模型,这在NLP中(特别是在语义和语法方面)通常是由熟练的注释者在耗时和昂贵的过程中生成的。其次,对大量数据的推断计算量很大,并且对环境有很强的影响。最后,深度学习模型对人类来说更难以解释。在深度学习模型中注入语言知识,尤其是“结构”知识可能会克服这些限制。要做到这一点,必须确定如何使深度学习模型意识到语言结构,并找到正确的结构偏差。一个(流行的)研究方向是用潜在变量对语法树等语言结构方面进行建模。从数据中学习这些模型需要特定的推理算法,通常与减少方差的技巧相结合。如果和当推断的模型不能提高性能,或者学习到的结构与语言学家的期望有很大的不同(往往是这种情况),那么总是有一个问题,这是否表明学习“正确”的结构失败,或者相反,假设的结构是否不能像模型所做的那样很好地解释数据。为了测试这些推理技术捕捉语言学家理论结构的严格能力,我们建议在“人工”数据集上评估这些算法。控制实际控制数据生成及其分布的结构应该能够测试这些算法学习给定类型结构的能力,而不需要考虑这些结构在建模真实语言数据和其他数据集特定问题方面的突出性。我们将特别强调推断(组合)语义解析的潜在语法树,并设计我们的人工数据集,以展示受语义学家工作启发的特定结构方面。
英文摘要
Deep learning is now ubiquitous in natural language processing (NLP). It has raised performances to new levels in a variety of tasks, including (but not limited to) : language modeling, syntactic and semantic parsing, named entity recognition, sentiment analysis and natural language understanding. At the heart of deep learning lies the idea of automatically learning and extracting features from the data that are useful to the task at end, thereby saving developers the need to handcraft all linguistic knowledge into computational models. Doing so counters some difficulties encountered by rule-based approaches. It notably offers a much easier handling of out-of-vocabulary items (leveraging pre-trained word embeddings), and a more expressive set of models with less commitments to (overly) strong assumptions (such as context-freeness hypotheses in approaches based on formal language theory, for example). These strong advantages come however with some downsides. First, very large amounts of data are generally required to infer good models, which in NLP (especially in semantics and syntax) is often produced by skilled annotators in a time consuming and expensive process. Second, inference on large amounts of data is computationally heavy and has a strong environmental impact. Finally, deep learning models are much more difficult for humans to interpret. Injecting linguistic knowledge, especially *structural* knowledge into deep learning models might overcome these limitations. To do this, one must determine how to make deep learning models aware of linguistic structure, and find the right kind of structural bias. A (popular) direction of research models aspects of linguistic structure like syntax trees with latent variables. Learning these models from data requires specific inference algorithms, often in combination with variance-reducing tricks. If and when the inferred model do not improve performance, or learned structures differ a lot from linguists' expectations (which tend to be the case), there is always a question of whether this indicates a failure to learn the 'correct' structures, or to the contrary whether the putative structures do not provide as good an explanation of the data as what the model is doing. We propose, in order to test the strict ability of these inference techniques to capture the kind of structures theorized by linguists, to evaluate these algorithms on *artificial* datasets. Controlling the structures actually governing the data generation and their distribution should enable testing the ability of these algorithms to learn a given type of structure, independently from concerns pertaining to the salience of these structures for modeling real linguistic data and other dataset-specific issues. We will put a particular emphasis on inferring latent syntactic trees for (compositional) semantic parsing, and design our artificial datasets to exhibit specific structural aspects inspired from the work of semanticists.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Assessing latent-structure models in natural language processing with artificial datasets
-
批准号:RGPIN-2021-03134
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.1万
-
财政年份:2021
-
负责人:Venant, Antoine
-
依托单位:
Assessing latent-structure models in natural language processing with artificial datasets
-
批准号:DGECR-2021-00145
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Venant, Antoine
-
依托单位:
国内基金
海外基金
基于LMP-1第五跨膜结构域为靶点治疗EB病毒诱导鼻咽癌的药物研发
-
批准号:21602216
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2016
-
负责人:王晓辉
-
依托单位: