课题基金 / 基金详情

Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains

Collaborative Research: RI: Medium: Expert-in-the-Loop Neural Summarization for Consequential Domains
合作研究:RI:中:结果领域的专家在环神经摘要
批准号:
2211955
负责人:
Zachary Lipton
金额:
$58.37万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2026-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
自动摘要方法旨在创建文本(例如新闻或科学文章)的缩短版本,但仍能准确地传达其要点。摘要方法提供了一种潜在的方法来抵消在许多领域普遍存在的“信息过载”问题。但很多关于自动摘要的研究主要集中在一种数据上:新闻文章。这并不是因为总结新闻文章被视为特别重要。更确切地说,这是由于有方便可用的大型数据集,可以用来“训练”机器学习模型来执行摘要。然而,由此产生的对假设可以访问大量“训练数据”以用于训练总结模型的方法的关注扭曲了研究优先级;很少有人研究自动摘要方法如何在医学或法律等重要但专门的领域中使用。在这些领域,人们不太可能有机会访问大量手工编写的摘要数据集。此外,这些领域的领域专家不太可能盲目地相信系统生成的摘要(他们也不应该)。这激发了对模型如何生成特定摘要的透明度的需求,以及允许专家更普遍地与模型交互的方法的需求。该项目旨在通过研究和扩展现代、预训练的神经总结模型的能力来解决这些问题,在这些领域和任务的背景下,人们没有明确的监督,并且对事实准确的总结有更高的需求。该项目将涉及对最先进的模型进行批判性评估,并对其进行微调,以便在有限监督下的医学等领域进行总结;一个特定的目标是根据模型输出的真实性来描述它们的行为。然后,我们的想法是扩展这些模型,通过主动学习方法,替代类型的监督(例如,专家“亮点”)和新的预训练目标,允许互动和有效的监督。最后,研究人员将设计提供更高透明度和可控性的架构;这将使用潜在变量汇总模型来完成,这将反过来允许人们检查哪些输入段通知了特定的输出。这将为最终用户(领域专家)验证模型输出提供一种自然的方法,并且它还将提供一种“调试”摘要系统的方法。希望这些技术创新将允许领域专家从自动摘要技术中获益。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Automatic summarization methods aim to create shortened versions of texts (for example, news or scientific articles) that still accurately communicate their main points. Summarization methods provide a potential means to counteract the problem of "information overload" which is prevalent across many areas. But much of the research on automatic summarization has focussed largely on just one type of data: news articles. This is not because summarizing news articles is seen as particularly important. Rather, it is a result of there being conveniently available large datasets that can be used to "train" machine learning models to perform summarization. However, the resultant focus on approaches that assume a setting in which one has access to large volumes of "training data" to use to train summarization models has warped research priorities; little work has been done on investigating how automatic summarization methods might be used in important but specialized domains such as medicine or law. In these kinds of areas one is unlikely to have access to a massive dataset of manually written summaries. Furthermore, domain experts in such areas are not likely to blindly trust a system-generated summary (nor should they). This motivates a need for transparency with respect to how the model generated a particular summary, and for approaches that permit the expert to interact with the model more generally. This project aims to address these issues by investigating and extending the capabilities of modern, pre-trained, neural summarization models in the context of domains and tasks in which one has limited explicit supervision, and where there is a heightened need for factually accurate summaries. The project will involve critically evaluating state-of-the-art models when fine-tuned for summarization in domains like medicine under limited supervision; a specific aim is to characterize their behavior with respect to the factuality of model outputs. The idea is then to extend these models to permit interactive and efficient supervision, via active learning methods, alternative types of supervision (e.g., expert "highlights"), and novel pre-training objectives. Finally, the investigators will design architectures that afford increased transparency and controllability; this will be accomplished using latent variable summarization models, which will in turn allow one to inspect which input segments informed particular outputs. This will provide a natural means for the end-user (domain expert) to verify model outputs, and it will also provide a means to "debug" summarization systems. The hope is that these technical innovations will allow domain experts to benefit from automated summarization technology.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)