课题基金 / 基金详情

Low-Resource Text Understanding

Low-Resource Text Understanding
低资源文本理解
批准号:
RGPIN-2021-03115
负责人:
Liu, Bang
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Liu, Bang的其他基金

相似基金

相关文献

中文摘要
翻译
自然语言理解(NLU)旨在阅读以自然语言形成的文本,确定每个元素的含义(例如,单词、句子、段落),并根据这些文本进行推理。它对各种应用程序至关重要,例如问答(QA)和对话系统。最先进的方法(例如,BERT,GPT-3)的NLU应用程序主要基于深度神经网络。然而,这些方法有三个主要缺点:i)它们是数据饥饿和计算密集型的。需要大量的训练数据和计算资源来训练这样的模型。用于问答的训练数据集的构建特别困难和耗时。(2)由于缺乏常识,他们经常在简单的问题上失败。(三)不可转让,缺乏解释性。该研究计划的目标是创建一个低资源,知识授权和逻辑增强的NLU系统,并将其应用于问答。它将有助于减少数据集构建中的人工工作,并使QA系统更具知识感知性,可解释性和可转移性。这些属性对于医学问题回答尤其重要,因为训练数据集非常难以构建,需要大量的领域知识来理解医学文本,并且向人类解释模型预测是必不可少的,因为医疗保健中的错误结果可能是昂贵和危险的。除了QA之外,本研究计划中将开发的技术对于NLU应用也至关重要。我们最近的论文发表在WWW 2020和SIGMOD 2020上,提出了从未标记的文本和Web用户搜索日志自动生成问题和Web规模本体构建的有效方法。我们将扩展我们以前的研究,并进一步探索以下方向:i)数据集构建。我们计划以完全受控的方式从未标记的文本和知识图中自动生成大规模和高质量的问答对,以减少人工劳动。我们还将评估数据质量,并学习有效的数据课程,以提高培训效率。(二)知识拓展。我们不是从头开始构建知识图,而是专注于本体扩展,用新发现的概念或实体扩展现有本体,以捕获世界上新兴的知识,并保持本体动态更新。生成的本体可以作为问题生成和问题回答的输入。(三)逻辑增强。我们将设计神经-符号问答模型,并将其与外部知识相结合,以提高问答系统的逻辑推理能力和可移植性。该计划将有助于加拿大在人工智能领域的领先地位,提高对自然语言的理解,并对医疗保健人工智能等现实世界的应用产生重大影响。
英文摘要
Natural language understanding (NLU) aims at reading texts formed in natural languages, determining the meaning of each element (e.g., words, sentences, paragraphs), and making inferences based on these texts. It is critical to various applications, such as question answering (QA) and dialogue systems. State-of-the-art methods (e.g., BERT, GPT-3) for NLU applications are mostly based on deep neural networks. However, these methods have three major drawbacks: i) They are data-hungry and computational-intensive. A lot of training data and computational resources are needed to train such models. The training datasets for question answering are especially difficult and time-consuming to build. ii) They often fail at simple questions due to a lack of commonsense knowledge. iii) They are not transferable and lack explainability. The goal of this research program is to create a low-resource, knowledge-empowered, and logic-enhanced NLU system, and apply it to question answering. It will help to reduce the human efforts in dataset construction, as well as make QA systems be more knowledge-aware, explainable, and transferable. Such properties are especially critical for medical question answering, where the training dataset is extremely hard to construct, a lot of domain knowledge is required to understand medical text, and giving explanations of the model predictions to humans is essential as an incorrect result in health-care can be costly and dangerous. The techniques that will be developed in this research program are also essential to NLU applications other than QA. Our recent papers, published at WWW 2020 and SIGMOD 2020, propose efficient approaches for automatic question generation and web-scale ontology construction from unlabeled text and web user search logs. We will extend our previous research and further explore the following directions: i) Dataset construction. We plan to automatically generate large-scale and high-quality question-answer pairs from unlabeled text and knowledge graphs in a fully controlled manner to reduce human efforts. We will also evaluate the data quality and learn an efficient data curriculum to improve training efficiency. ii) Knowledge expansion. Instead of constructing a knowledge graph from scratch, we focus on ontology expansion to expand an existing ontology with newly discovered concepts or entities to capture the emerging knowledge in the world and keep the ontology dynamically updated. The generated ontology can serve as an input for both question generation and question answering. iii) Logic enhancement. We will design neural-symbolic QA models and integrate them with external knowledge so that we can improve both the logical reasoning ability and the transferability of QA systems. This program will contribute to Canada's lead in artificial intelligence, improve the understanding of natural language, and has a significant impact on real-world applications such as AI for health-care.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Low-Resource Text Understanding
  • 批准号:
    DGECR-2021-00316
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2021
  • 负责人:
    Liu, Bang
  • 依托单位:
Low-Resource Text Understanding
  • 批准号:
    RGPIN-2021-03115
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Liu, Bang
  • 依托单位:
海外基金