Low-Resource Text Understanding
Low-Resource Text Understanding
批准号:
RGPIN-2021-03115
负责人:
Liu, Bang
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
自然语言理解的目的是阅读自然语言形成的文本,确定每个元素(如单词、句子、段落)的含义,并在这些文本的基础上进行推理。它对于问答(QA)和对话系统等各种应用都是至关重要的。用于NLU应用的最新方法(例如,BERT、GPT-3)大多基于深度神经网络。然而,这些方法有三个主要缺点:i)它们需要大量的数据和大量的计算。训练这样的模型需要大量的训练数据和计算资源。建立问题回答的训练数据集尤其困难和耗时。Ii)由于缺乏常识性知识,他们经常在简单的问题上失败。三)不可转让,缺乏可解释性。本研究项目的目标是创建一个低资源、知识赋能、逻辑增强的自然语言理解系统,并将其应用于问答系统。这将有助于减少人工构建数据集的工作量,并使QA系统更具知识意识、可解释性和可移植性。这样的属性对于医学问题回答尤其关键,其中训练数据集极其难以构建,需要大量领域知识来理解医学文本,并且向人类解释模型预测是必不可少的,因为医疗保健中的错误结果可能是昂贵和危险的。这项研究计划中将开发的技术对于除QA之外的NLU应用程序也是必不可少的。我们最近在WWW 2020和SIGMOD 2020上发表的论文提出了从未标记文本和Web用户搜索日志自动生成问题和构建Web规模本体的有效方法。我们将在已有研究的基础上,进一步探索以下方向:i)数据集构建。我们计划以完全受控的方式从未标记的文本和知识图中自动生成大规模和高质量的问答对,以减少人工工作。我们还将评估数据质量,学习高效的数据课程,以提高培训效率。二、知识扩张。我们不是从头开始构建知识图谱,而是专注于本体扩展,用新发现的概念或实体扩展现有的本体,以捕获世界上正在出现的知识,并保持本体的动态更新。生成的本体可以用作问题生成和问题回答的输入。(三)逻辑强化。我们将设计神经符号问答模型,并将其与外部知识相结合,从而提高问答系统的逻辑推理能力和可移植性。该项目将有助于加拿大在人工智能领域的领先地位,提高对自然语言的理解,并对医疗保健的人工智能等现实世界应用产生重大影响。
英文摘要
Natural language understanding (NLU) aims at reading texts formed in natural languages, determining the meaning of each element (e.g., words, sentences, paragraphs), and making inferences based on these texts. It is critical to various applications, such as question answering (QA) and dialogue systems. State-of-the-art methods (e.g., BERT, GPT-3) for NLU applications are mostly based on deep neural networks. However, these methods have three major drawbacks: i) They are data-hungry and computational-intensive. A lot of training data and computational resources are needed to train such models. The training datasets for question answering are especially difficult and time-consuming to build. ii) They often fail at simple questions due to a lack of commonsense knowledge. iii) They are not transferable and lack explainability. The goal of this research program is to create a low-resource, knowledge-empowered, and logic-enhanced NLU system, and apply it to question answering. It will help to reduce the human efforts in dataset construction, as well as make QA systems be more knowledge-aware, explainable, and transferable. Such properties are especially critical for medical question answering, where the training dataset is extremely hard to construct, a lot of domain knowledge is required to understand medical text, and giving explanations of the model predictions to humans is essential as an incorrect result in health-care can be costly and dangerous. The techniques that will be developed in this research program are also essential to NLU applications other than QA. Our recent papers, published at WWW 2020 and SIGMOD 2020, propose efficient approaches for automatic question generation and web-scale ontology construction from unlabeled text and web user search logs. We will extend our previous research and further explore the following directions: i) Dataset construction. We plan to automatically generate large-scale and high-quality question-answer pairs from unlabeled text and knowledge graphs in a fully controlled manner to reduce human efforts. We will also evaluate the data quality and learn an efficient data curriculum to improve training efficiency. ii) Knowledge expansion. Instead of constructing a knowledge graph from scratch, we focus on ontology expansion to expand an existing ontology with newly discovered concepts or entities to capture the emerging knowledge in the world and keep the ontology dynamically updated. The generated ontology can serve as an input for both question generation and question answering. iii) Logic enhancement. We will design neural-symbolic QA models and integrate them with external knowledge so that we can improve both the logical reasoning ability and the transferability of QA systems. This program will contribute to Canada's lead in artificial intelligence, improve the understanding of natural language, and has a significant impact on real-world applications such as AI for health-care.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Low-Resource Text Understanding
-
批准号:RGPIN-2021-03115
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.48万
-
财政年份:2022
-
负责人:Liu, Bang
-
依托单位:
Low-Resource Text Understanding
-
批准号:DGECR-2021-00316
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Liu, Bang
-
依托单位:
海外基金