Automated Extraction of Fields of Interest from Unstructured Documents
Automated Extraction of Fields of Interest from Unstructured Documents
批准号:
571403-2021
负责人:
Ng, RaymondT
金额:
$6.47万
依托单位国家:
加拿大
项目类别:
Alliance Grants
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
人类经验的书面知识是巨大而丰富的。然而,很多这方面的知识是未知的和难以接近的。自然语言处理(NLP)是通过自动理解文本来解锁这些有价值信息的一种途径。这种潜力正在随着最近出现的基础或大型语言模型的出现而实现,这些模型将自我监督的深度神经网络的力量与不可思议的大而广泛的数据集相结合。这些模型的一个特点是“迁移学习”的概念,即模型被训练完成一项任务,并可以进行微调以完成另一项相关但不同的任务。在本提案中,该团队将重点关注NLP中大型语言模型(例如BERT, GPT-3)的最新进展,以从非结构化文档中提取有价值的信息。尽管在非结构化文档中缺乏内聚和组织,但它们通常带有嵌入式结构化数据。将有用的、结构化的信息表示为表格是科学、金融、卫生等许多领域中用于信息检索和其他任务的一种常用而有效的方法。然而,手动从文档中提取结构化数据通常需要花费大量的时间和人力,这就需要一个自动化流程的系统。具体来说,该提案将利用BERT和T5模型的突破性进展来开发从文档中自动提取表的工具。也就是说,给定一组“感兴趣的字段”(FoI),任务是自动提取并填充一个表,并将FoI作为列标题。为了说明我们的工具具有广泛的适用性,我们在两个非常不同的应用领域展示了我们的算法的有效性:(1)乳腺癌护理;(2)矿产开采。在提取这些表之后,数据可以用于各种“下游”分析任务,例如问题回答和预测建模。
英文摘要
The written knowledge of human experience is immense and rich. However, much of this knowledge is unbeknownst and inaccessible. Natural language processing (NLP) represents one avenue to unlock this valuable information by understanding text automatically. This potential is being realized with the recent emergence of foundation or large language models that combine the power of self-supervised deep neural networks with unfathomably large and broad data sets. One hallmark of these models is the concept of "transfer learning," whereby the model is trained to do one task and can be fine-tuned to complete a related but different task. In this proposal, the team will focus on recent advances with large language models (e.g., BERT, GPT-3) in NLP to extract valuable information from unstructured documents. Despite the lack of cohesion and organization in unstructured documents, they often come with embedded structured data. Representing useful and structured information as tables is a common and effective method in scientific, financial, health and many domains for information retrieval and other tasks. However, manual extraction of structured data from documents typically costs tremendous time and labour, motivating the need for a system for automating the process. Specifically, this proposal will leverage groundbreaking advances by BERT and T5 models to develop tools for automated table extraction from documents. That is, given a collection of "fields of interest" (FoI), the task is to automatically extract and populate a table with the FoIs as column headers. To illustrate that our tools are general with broad applicability, we demonstrate the effectiveness of our algorithms with two very different application domains: (1) breast cancer care; and (2) mineral mining. After such tables have been extracted, the data can be used for various "downstream" analytics tasks such as question answering and predictive modeling.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金