Automated Extraction of Fields of Interest from Unstructured Documents
Automated Extraction of Fields of Interest from Unstructured Documents
批准号:
571403-2021
负责人:
Ng, RaymondT
金额:
$6.47万
依托单位国家:
加拿大
项目类别:
Alliance Grants
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The written knowledge of human experience is immense and rich. However, much of this knowledge is unbeknownst and inaccessible. Natural language processing (NLP) represents one avenue to unlock this valuable information by understanding text automatically. This potential is being realized with the recent emergence of foundation or large language models that combine the power of self-supervised deep neural networks with unfathomably large and broad data sets. One hallmark of these models is the concept of "transfer learning," whereby the model is trained to do one task and can be fine-tuned to complete a related but different task. In this proposal, the team will focus on recent advances with large language models (e.g., BERT, GPT-3) in NLP to extract valuable information from unstructured documents. Despite the lack of cohesion and organization in unstructured documents, they often come with embedded structured data. Representing useful and structured information as tables is a common and effective method in scientific, financial, health and many domains for information retrieval and other tasks. However, manual extraction of structured data from documents typically costs tremendous time and labour, motivating the need for a system for automating the process. Specifically, this proposal will leverage groundbreaking advances by BERT and T5 models to develop tools for automated table extraction from documents. That is, given a collection of "fields of interest" (FoI), the task is to automatically extract and populate a table with the FoIs as column headers. To illustrate that our tools are general with broad applicability, we demonstrate the effectiveness of our algorithms with two very different application domains: (1) breast cancer care; and (2) mineral mining. After such tables have been extracted, the data can be used for various "downstream" analytics tasks such as question answering and predictive modeling.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金