BSF: 2016257: Building Models for Reading Comprehension in Specialized Domains from Scratch
BSF: 2016257: Building Models for Reading Comprehension in Specialized Domains from Scratch
批准号:
1737230
负责人:
Vivek Srikumar
金额:
$3.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2020-08-31
中文摘要
机器学习算法越来越多地允许人们搜索、构建和访问每天在每个可能的领域中创建的文本信息。在有大量注释数据的领域,可以应用监督学习算法,用于文本理解的算法在构建文本和提供自然语言接口方面取得了成功。然而,当在一个几乎没有数据的新领域中构建系统时,数据收集和注释数据可能会非常昂贵。该项目探索了一种协议,用于开发文本理解系统,该系统可以读取文本并在特定领域(如生物学或历史)提供自然语言接口——这可以允许专业社区对文本中锁定的数据进行数字访问。该项目还培训学生,作为国际合作的一部分——该奖项支持美国研究人员在美国-以色列两国科学基金会资助的一个项目中进行合作。该项目包括数据收集和模型训练,并考虑两者之间的相互作用。为了取代专家的注释,它使用众包工作人员在一个迭代过程中开始训练一个几乎没有数据的结构化预测器。它创建了一个交互式框架,用户可以在其中提出问题并验证候选答案,这些答案随后用于对系统进行再培训。它的目标是在多个领域进行联合训练,并使用领域自适应方法将知识从一个领域转移到另一个领域。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning algorithms are increasingly allowing people to search, structure, and access the textual information created daily in every possible domain. In areas with abundant annotated data, where supervised learning algorithms can be applied, algorithms for text understanding have had success in structuring text and providing natural language interfaces. When building a system in a new domain for which there is little to no data, however, data collection and annotation data can be prohibitively expensive. This project explores a protocol for developing text understanding systems that read text and provide a natural language interface in a particular domain (such as biology or history) -- this can allow specialized communities to have digital access to data that is otherwise locked in text. The project also trains students as part of an international collaboration -- this award supports travel of the US-based researchers to collaborate in a project funded by the US-Israel Binational Science Foundation. The project encompasses both data collection and model training, and considers the interaction between the two. To replace expert annotations it uses crowdsourcing workers in an iterative procedure that starts training a structured predictor from almost no data. It creates an interactive framework in which users ask questions and verify candidate answers that are later used to retrain the system. It aims to jointly train over multiple domains, and use domain adaptation methods to transfer knowledge from one domain to another. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.18653/v1/k19-1042
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
作者:
[Omri Koshorek;Gabriel Stanovsky;Yichu Zhou;Vivek Srikumar;Jonathan Berant]
通讯作者:
Omri Koshorek;Gabriel Stanovsky;Yichu Zhou;Vivek Srikumar;Jonathan Berant
III: Small: Collaborative Research: Scrutable and Explainable Information Retrieval with Model Intrinsic and Agnostic Approaches
-
批准号:2007398
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2020
-
负责人:Vivek Srikumar
-
依托单位:
Technology Facilitated Training for Mental Health Counseling
-
批准号:1822877
-
项目类别:Standard Grant
-
资助金额:$74.98万
-
财政年份:2018
-
负责人:Vivek Srikumar
-
依托单位: