Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
批准号:
2211954
负责人:
Byron Wallace
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2026-06-30
中文摘要
自动摘要方法旨在创建文本的缩短版本(例如,新闻或科学文章),仍然准确地传达其主要观点。摘要方法提供了一种潜在的手段,以抵消在许多领域普遍存在的“信息过载”问题。但是,许多关于自动摘要的研究主要集中在一种类型的数据上:新闻文章。这并不是因为总结新闻文章被认为是特别重要的。相反,这是由于存在方便可用的大型数据集,可用于“训练”机器学习模型以执行摘要。然而,由此产生的重点是假设一个设置,其中一个可以访问大量的“训练数据”来训练摘要模型的方法已经扭曲了研究的优先事项;很少有工作已经做了调查如何自动摘要方法可能会被用于重要的,但专门的领域,如医学或法律。在这些领域,人们不太可能获得大量的手工编写摘要的数据集。此外,这些领域的领域专家不太可能盲目相信系统生成的摘要(他们也不应该)。这激发了对模型如何生成特定摘要的透明度的需求,以及允许专家更普遍地与模型交互的方法。该项目旨在通过调查和扩展现代,预先训练的神经摘要模型的能力来解决这些问题,这些模型适用于那些明确监督有限的领域和任务,并且对事实准确的摘要有更高的需求。该项目将涉及在有限监督下对医学等领域的总结进行微调时,对最先进的模型进行批判性评估;一个具体的目标是描述它们在模型输出的真实性方面的行为。然后,我们的想法是扩展这些模型,以允许交互式和有效的监督,通过主动学习方法,替代类型的监督(例如,专家“亮点”)和新的预培训目标。最后,研究人员将设计提供更高透明度和可控性的架构;这将使用潜变量汇总模型来实现,这反过来又允许人们检查哪些输入段通知了特定的输出。这将为最终用户(领域专家)提供一种验证模型输出的自然方法,也将提供一种“调试”摘要系统的方法。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automatic summarization methods aim to create shortened versions of texts (for example, news or scientific articles) that still accurately communicate their main points. Summarization methods provide a potential means to counteract the problem of “information overload” which is prevalent across many areas. But much of the research on automatic summarization has focussed largely on just one type of data: news articles. This is not because summarizing news articles is seen as particularly important. Rather, it is a result of there being conveniently available large datasets that can be used to “train” machine learning models to perform summarization. However, the resultant focus on approaches that assume a setting in which one has access to large volumes of “training data” to use to train summarization models has warped research priorities; little work has been done on investigating how automatic summarization methods might be used in important but specialized domains such as medicine or law. In these kinds of areas one is unlikely to have access to a massive dataset of manually written summaries. Furthermore, domain experts in such areas are not likely to blindly trust a system-generated summary (nor should they). This motivates a need for transparency with respect to how the model generated a particular summary, and for approaches that permit the expert to interact with the model more generally. This project aims to address these issues by investigating and extending the capabilities of modern, pre-trained, neural summarization models in the context of domains and tasks in which one has limited explicit supervision, and where there is a heightened need for factually accurate summaries. The project will involve critically evaluating state-of-the-art models when fine-tuned for summarization in domains like medicine under limited supervision; a specific aim is to characterize their behavior with respect to the factuality of model outputs. The idea is then to extend these models to permit interactive and efficient supervision, via active learning methods, alternative types of supervision (e.g., expert “highlights”), and novel pre-training objectives. Finally, the investigators will design architectures that afford increased transparency and controllability; this will be accomplished using latent variable summarization models, which will in turn allow one to inspect which input segments informed particular outputs. This will provide a natural means for the end-user (domain expert) to verify model outputs, and it will also provide a means to “debug” summarization systems. The hope is that these technical innovations will allow domain experts to benefit from automated summarization technology.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Automatically Summarizing Evidence from Clinical Trials: A Prototype Highlighting Current Challenges
DOI:
10.48550/arxiv.2303.05392
发表时间:
2023-03
期刊:
Proceedings of the conference. Association for Computational Linguistics. Meeting
影响因子:
--
作者:
[S. Ramprasad;Denis Jered McInerney;Iain J. Marshal;Byron Wallace]
通讯作者:
S. Ramprasad;Denis Jered McInerney;Iain J. Marshal;Byron Wallace
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Lucy Lu Wang;Jay DeYoung;Byron Wallace]
通讯作者:
Lucy Lu Wang;Jay DeYoung;Byron Wallace
RI: Medium: Learning Disentangled Representations for Text to Aid Interpretability and Transfer
-
批准号:1901117
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2019
-
负责人:Byron Wallace
-
依托单位:
CAREER: Structured Scientific Evidence Extraction: Models and Corpora
-
批准号:1750978
-
项目类别:Continuing Grant
-
资助金额:$54.99万
-
财政年份:2018
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1520781
-
项目类别:Standard Grant
-
资助金额:$25.59万
-
财政年份:2014
-
负责人:Byron Wallace
-
依托单位:
Collaborative research: ABI Development: Making Advanced Statistical Tools Accessible for Quantitative Research Synthesis and Discovery in Ecology and Evolutionary Biology
-
批准号:1262442
-
项目类别:Standard Grant
-
资助金额:$50.27万
-
财政年份:2013
-
负责人:Byron Wallace
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: