课题基金 / 基金详情

QASciInf: Question Answering for Scientific Information

QASciInf: Question Answering for Scientific Information
QASciInf:科学信息问答
批准号:
252295018
负责人:
Professorin Dr. Iryna Gurevych
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Professorin Dr. Iryna Gurevych的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的几十年里,发表的科学论文的数量呈指数级增长。这使得研究人员很难找到并从所有相关工作中受益。在这个项目中,我们解决了这个问题,并提出了一套新颖的研究技术来对科学信息进行问答(QA)。科学领域的独特挑战需要迄今为止QA研究中尚未探索的新方法。特别是,科学信息的QA系统需要(a)考虑来自异构来源的信息,(b)包括更好的上下文感知方法来处理科学文章所代表的长上下文,以及(c)对表的内容进行推理,以根据数据生成答案。为了实现这一方向的研究,我们构建了两个新的数据集,用于(1)对科学文章、表格数据和网络讨论的文本进行混合问答,以及(2)通过对表格内容的推理生成信息表描述。与现有的QA数据集相比,科学QA并不局限于解决文章文本的问答对。有些问题只能通过科学表格的推理来回答,有些问题可以通过网络上的相关讨论来回答。因此,基于这些数据集,我们提出了可以从网络讨论中选择相关内容的方法,同时结合来自科学文章和检索讨论的丰富上下文信息。此外,我们还研究了能够对复杂科学表格进行推理的新型文本生成模型。由于推理感知的表到文本生成需要大量的训练数据,我们提出了一种新的方法,通过使用弱监督和半监督训练技术自动扩展训练数据来训练可泛化的表到文本模型。最后,我们将模型整合到混合QA原型中,并在用户研究中对科学文献进行评估。
英文摘要
The number of published scientific articles has grown exponentially in the last few decades. This makes it hardly possible for researchers to find and benefit from all relevant works. In this project, we address this problem and propose a set of novel research techniques to perform question answering (QA) over scientific information. The unique challenges of the scientific domain require novel approaches that are, to date, unexplored in QA research. In particular, a QA system for scientific information needs to (a) consider information from heterogeneous sources, (b) include better context-aware methods to process the long context that is represented by scientific articles, and (c) reason over the content of tables to generate answers based on the data. To enable research in this direction, we construct two novel datasets for (1) hybrid question answering over the text of scientific articles, table data, and discussions on the web, and (2) generating informative table descriptions through reasoning over the table content. In contrast to existing QA datasets, scientific QA is not limited to question-answer pairs that address the text of articles. Some questions can only be answered by reasoning over scientific tables and some can be answered by using related discussions on the web. Thus, based on these datasets, we propose approaches that can select relevant content from discussions on the web, while incorporating rich contextual information from the scientific article and the retrieved discussions. In addition, we research novel text generation models that are capable of reasoning over complex scientific tables. Because reasoning-aware table-to-text generation requires a considerable amount of training data, we propose novel methods to train generalizable table-to-text models by automatically expanding the training data with weakly supervised and semi-supervised training techniques. Finally, we consolidate our models in a prototype for hybrid QA over scientific literature which we evaluate in a user study.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Open Argument Mining
Argumentation Analysis for the Web
Feature-based Visualization and Analysis of Natural Language Documents
Integrating Collaborative and Linguistic Resources for Word Sense Disambiguation and Semantic Role Labeling (InCoRe)
海外基金