Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
批准号:
2211954
负责人:
Byron Wallace
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2026-06-30
中文摘要
自动摘要方法旨在创建仍然准确传达要点的文本缩短版本(例如新闻或科学文章)。摘要方法提供了一种潜在的方法来解决许多领域普遍存在的“信息过载”问题。但许多有关自动摘要的研究主要集中在一种类型的数据上:新闻文章。这并不是因为总结新闻文章被认为特别重要。相反,这是因为存在方便可用的大型数据集,可用于“训练”机器学习模型以执行摘要。然而,由此产生的对假设一种环境的方法的关注,在这种环境中,人们可以访问大量的“训练数据”来训练摘要模型,这扭曲了研究的重点;关于如何在医学或法律等重要但专业的领域中使用自动摘要方法的研究工作很少。在这些领域,人们不太可能访问大量手动编写的摘要数据集。此外,这些领域的领域专家不太可能(也不应该)盲目地相信系统生成的摘要。这就激发了对模型如何生成特定摘要的透明度的需求,以及允许专家更广泛地与模型交互的方法的需求。该项目旨在通过研究和扩展现代预训练神经摘要模型的能力来解决这些问题,这些模型在明确监督有限且非常需要准确摘要的领域和任务中。该项目将涉及批判性地评估最先进的模型,并在有限监督下对医学等领域的总结进行微调;一个具体目标是根据模型输出的真实性来描述他们的行为。然后,我们的想法是通过主动学习方法、替代类型的监督(例如专家“亮点”)和新颖的预训练目标来扩展这些模型,以实现交互式和高效的监督。最后,研究人员将设计能够提高透明度和可控性的架构;这将使用潜在变量汇总模型来完成,这反过来又允许人们检查哪些输入段通知了特定的输出。这将为最终用户(领域专家)提供一种验证模型输出的自然方法,并且还将提供一种“调试”摘要系统的方法。希望这些技术创新将使领域专家能够从自动摘要技术中受益。该奖项反映了 NSF 的法定使命,并通过使用基金会的智力优点和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automatic summarization methods aim to create shortened versions of texts (for example, news or scientific articles) that still accurately communicate their main points. Summarization methods provide a potential means to counteract the problem of “information overload” which is prevalent across many areas. But much of the research on automatic summarization has focussed largely on just one type of data: news articles. This is not because summarizing news articles is seen as particularly important. Rather, it is a result of there being conveniently available large datasets that can be used to “train” machine learning models to perform summarization. However, the resultant focus on approaches that assume a setting in which one has access to large volumes of “training data” to use to train summarization models has warped research priorities; little work has been done on investigating how automatic summarization methods might be used in important but specialized domains such as medicine or law. In these kinds of areas one is unlikely to have access to a massive dataset of manually written summaries. Furthermore, domain experts in such areas are not likely to blindly trust a system-generated summary (nor should they). This motivates a need for transparency with respect to how the model generated a particular summary, and for approaches that permit the expert to interact with the model more generally. This project aims to address these issues by investigating and extending the capabilities of modern, pre-trained, neural summarization models in the context of domains and tasks in which one has limited explicit supervision, and where there is a heightened need for factually accurate summaries. The project will involve critically evaluating state-of-the-art models when fine-tuned for summarization in domains like medicine under limited supervision; a specific aim is to characterize their behavior with respect to the factuality of model outputs. The idea is then to extend these models to permit interactive and efficient supervision, via active learning methods, alternative types of supervision (e.g., expert “highlights”), and novel pre-training objectives. Finally, the investigators will design architectures that afford increased transparency and controllability; this will be accomplished using latent variable summarization models, which will in turn allow one to inspect which input segments informed particular outputs. This will provide a natural means for the end-user (domain expert) to verify model outputs, and it will also provide a means to “debug” summarization systems. The hope is that these technical innovations will allow domain experts to benefit from automated summarization technology.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Automatically Summarizing Evidence from Clinical Trials: A Prototype Highlighting Current Challenges
DOI:
10.48550/arxiv.2303.05392
发表时间:
2023-03
期刊:
Proceedings of the conference. Association for Computational Linguistics. Meeting
影响因子:
--
作者:
[S. Ramprasad;Denis Jered McInerney;Iain J. Marshal;Byron Wallace]
通讯作者:
S. Ramprasad;Denis Jered McInerney;Iain J. Marshal;Byron Wallace
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Lucy Lu Wang;Jay DeYoung;Byron Wallace]
通讯作者:
Lucy Lu Wang;Jay DeYoung;Byron Wallace
RI: Medium: Learning Disentangled Representations for Text to Aid Interpretability and Transfer
-
批准号:1901117
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2019
-
负责人:Byron Wallace
-
依托单位:
CAREER: Structured Scientific Evidence Extraction: Models and Corpora
-
批准号:1750978
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2018
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1520781
-
项目类别:Standard Grant
-
资助金额:$25.59万
-
财政年份:2014
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1262442
-
项目类别:Standard Grant
-
资助金额:$50.27万
-
财政年份:2013
-
负责人:Byron Wallace
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: