Low-Resource Text Understanding
Low-Resource Text Understanding
批准号:
RGPIN-2021-03115
负责人:
Liu, Bang
金额:
$2.48万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Natural language understanding (NLU) aims at reading texts formed in natural languages, determining the meaning of each element (e.g., words, sentences, paragraphs), and making inferences based on these texts. It is critical to various applications, such as question answering (QA) and dialogue systems. State-of-the-art methods (e.g., BERT, GPT-3) for NLU applications are mostly based on deep neural networks. However, these methods have three major drawbacks: i) They are data-hungry and computational-intensive. A lot of training data and computational resources are needed to train such models. The training datasets for question answering are especially difficult and time-consuming to build. ii) They often fail at simple questions due to a lack of commonsense knowledge. iii) They are not transferable and lack explainability. The goal of this research program is to create a low-resource, knowledge-empowered, and logic-enhanced NLU system, and apply it to question answering. It will help to reduce the human efforts in dataset construction, as well as make QA systems be more knowledge-aware, explainable, and transferable. Such properties are especially critical for medical question answering, where the training dataset is extremely hard to construct, a lot of domain knowledge is required to understand medical text, and giving explanations of the model predictions to humans is essential as an incorrect result in health-care can be costly and dangerous. The techniques that will be developed in this research program are also essential to NLU applications other than QA. Our recent papers, published at WWW 2020 and SIGMOD 2020, propose efficient approaches for automatic question generation and web-scale ontology construction from unlabeled text and web user search logs. We will extend our previous research and further explore the following directions: i) Dataset construction. We plan to automatically generate large-scale and high-quality question-answer pairs from unlabeled text and knowledge graphs in a fully controlled manner to reduce human efforts. We will also evaluate the data quality and learn an efficient data curriculum to improve training efficiency. ii) Knowledge expansion. Instead of constructing a knowledge graph from scratch, we focus on ontology expansion to expand an existing ontology with newly discovered concepts or entities to capture the emerging knowledge in the world and keep the ontology dynamically updated. The generated ontology can serve as an input for both question generation and question answering. iii) Logic enhancement. We will design neural-symbolic QA models and integrate them with external knowledge so that we can improve both the logical reasoning ability and the transferability of QA systems. This program will contribute to Canada's lead in artificial intelligence, improve the understanding of natural language, and has a significant impact on real-world applications such as AI for health-care.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Low-Resource Text Understanding
-
批准号:RGPIN-2021-03115
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.48万
-
财政年份:2022
-
负责人:Liu, Bang
-
依托单位:
Low-Resource Text Understanding
-
批准号:DGECR-2021-00316
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Liu, Bang
-
依托单位:
海外基金