III: Small: Matching and Ranking via Proximity Graphs: Applications to Question Answering and Beyond
III: Small: Matching and Ranking via Proximity Graphs: Applications to Question Answering and Beyond
批准号:
1618159
负责人:
Eric Nyberg
金额:
$49.85万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-08-31
中文摘要
该项目将探索经典的基于术语的全文搜索的新替代方案,这是使用最广泛的计算机算法之一。当前的全文搜索方法在很大程度上依赖于记住哪些单词和短语出现在哪些文本文档中。相反,拟议的研究将通过使用更一般的相似性搜索方法来检查偏离这一研究得很好的路径的方法。在这样做的过程中,拟议的研究将追求以下两个目标:(1)缓解现有方法的局限性,例如查询和文档中出现的单词之间的不匹配;(2)开发允许数据科学家和检索算法设计者有效分工的方法。后者将使数据科学家能够专注于开发有效的相似性模型,而不必太担心低级别的性能问题,而检索算法的设计者和软件工程师将能够专注于开发更高效和/或可扩展的方法,而对结果质量的担忧更少。拟议的研究将调查至少两种情况,即基于术语的全文搜索被更通用的高精度k最近邻(k-NN)搜索所取代。在第一个场景中,它将开发一个超越纯粹词汇匹配的相似性函数,并考虑分布相似性、从平行(单语言)语料库学习的相似性等。在这种情况下,相似性函数将作为黑盒函数与通用相似性搜索引擎相结合,作为非度量空间图书馆(NMSLIB)的一部分实现。我们将探索几种搜索算法。其中一种搜索方法将依赖于构建邻近图(也称为邻域图),其中节点是对象,相似的节点通过边连接。在第二个场景中,拟议的研究将在超项上构建伪倒排文件。超项是出现在小尺寸滑动窗口内的单词的(密集或稀疏)向量表示。超词形成可以使用接近图(或任何其他有效的k-NN搜索方法)来索引的伪词汇表。在查询时,将从查询中提取超级词语,并与伪词汇表进行匹配,以获得k个最近的超级词语(以及它们出现的文档)。这种方法将包括术语邻近度和术语相似度(后者将使该方法受到词汇不匹配的影响较小)。由于初步实验表明,对于手头的任务,接近图不够准确和有效,因此拟议的研究还将尝试开发更好的接近图方法变体。如果这种改进失败了,还将探索替代的搜索方法。实验的见解、算法的改进和新的具有挑战性的数据集(由拟议的工作产生)将推动k-NN搜索的技术水平,这是另一种广泛使用的方法。这反过来将有利于其他各种NLP任务,如分类、基于词典的实体检测和第一故事检测,这些任务都严重依赖于k-NN搜索。更多项目信息将在项目网站上提供:http://www.lti.cs.cmu.edu/PGraph
英文摘要
This project will explore novel alternatives to a classic term-based full-text search, which is one of the most widely used computer algorithms. The current full-text search approaches heavily rely on memorizing which words and phrases appear in which text documents. The proposed research, in contrast, will examine methods that deviate from this well-studied path by using more generic similarity search methods. In doing so, the proposed research will pursue the following two objectives: (1) mitigating limitations of the existing approaches such as the mismatch between words that appear in queries and documents; and (2) developing approaches that permit an efficient separation of labor between data scientists and designers of retrieval algorithms. The latter would allow data scientists to focus on development of effective similarity models without worrying too much about low-level performance issues, while designers of retrieval algorithms and software engineers will be able to focus on development of more efficient and/or scalable approaches having fewer concerns about quality of results.The proposed research will investigate at least two scenarios where a term-based full-text search is replaced with a more generic high-accuracy k-nearest (k-NN) neighbor search. In the first scenario, it will develop a similarity function that goes beyond pure lexical matching and takes into account distributional similarity, similarity learned from a parallel (monolingual) corpus, and so on. In this scenario, the similarity function will be used as a black-box function coupled with a generic similarity search engine, implemented as a part of the Non-Metric Space Library (NMSLIB). Several search algorithms will be explored. One of the search approaches will rely on building a proximity graph (also known as a neighborhood graph), where nodes are objects and similar nodes are connected by edges. In the second scenario, the proposed research will build a pseudo inverted file over super terms. Super terms are (dense or sparse) vectorial representations of words appearing within a sliding window of small size. The super terms form a pseudo-vocabulary that can be indexed using a proximity graph (or any other efficient k-NN search method). At query time, the super terms will be extracted from the query and matched against the pseudo-vocabulary to obtain k nearest super terms (as well as documents where they occur). This approach will incorporate term proximity and term similarity (the latter will make the approach less affected by the vocabulary mismatch). Because preliminary experiments demonstrated that proximity graphs are not sufficiently accurate and efficient for the task in hand, the proposed research will also attempt to develop better variants of the proximity graphs methods. Should such an improvement fail, alternative search methods will also be explored. Experimental insights, algorithmic improvements, and new challenging datasets (resulting from the proposed work) will advance the state of the art in k-NN search, which is another widely used method. This, in turn, will benefit a variety of other NLP tasks such as classification, dictionary-based entity detection, and first story detection, which all heavily relying on the k-NN search. Additional project information will be made available at the project website: http://www.lti.cs.cmu.edu/PGraph
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: