Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
批准号:
RGPIN-2015-03807
负责人:
Huang, Jimmy
金额:
$3.13万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
使用谷歌很容易。使用谷歌查找信息则是另一回事。几十年来,信息检索(IR)取得了长足的进步。然而,IR远未成为一个解决的问题,仍然存在许多挑战。首先,大多数Web搜索引擎将简短的文本查询作为输入,并输出排序的文档列表。检索决策主要基于当前的查询和文档集合。给定查询的结果通常是相同的,与用户或用户发出请求的上下文无关。第二,信息交流是一个互动的过程。在当前以文档为中心的检索范式下,交互式检索被视为一系列独立的简单检索决策步骤。然而,人们已经注意到,对任务感知用户会话的分析提供了对用户查询行为的有用洞察,其中任务感知用户会话包含由用户提交以满足信息需求的一系列请求。第三,目前的信息检索系统大多使用关键字来查询和索引文档。然而,这种传统的基于关键字的信息检索模型对用户信息需求的理解提供的语义上下文很少,这会导致搜索中查询和文档的不匹配。理想情况下,人们希望看到查询和文档相互匹配,如果它们是主题相关的话。因此,需要根据用户的信息需求和用户对文档的理解将语义上下文整合到信息检索系统中,以提高信息检索的性能。第四,组织现在越来越多地处理PB级的数据收集。在处理大数据时,情景敏感型和任务感知型方法变得更具挑战性。因此,提出能够有效、高效地处理大数据的新算法和模型,并在大数据场景中实施上下文敏感和任务感知的方法非常重要。*在数据以非常快的速度增长的世界中,对更准确、更有效地搜索和分析大数据以发现有用信息的需求非常巨大。这个研究项目解决了从大文本数据中搜索和发现有用信息的问题。这项研究的长期目标是克服现有信息检索方法的局限性,正式开发一种新的检索范式,称为上下文敏感和任务感知的大数据信息搜索。特别是,(1)我们将开发一个新的理论检索框架,以捕获丰富的用户信息并提供个性化的搜索结果;(2)我们将为上下文敏感信息检索开发新的基于任务的检索方法,以优化整个检索会话的长期检索效用;(3)我们将开发新的自动分析和搜索大数据的模型,以高效地提取知识和执行语义匹配。**
英文摘要
Using Google is easy. Finding information using Google is another matter. Over the decades, significant progress has been made in Information Retrieval (IR). However, IR is far from a solved problem and many challenges remain. First, most Web search engines take a short text query as input and output a ranked list of documents. The retrieval decision is made primarily based on the current query and document collection. The results for a given query are usually identical, independent of the user or the context in which the user made the request. Second, IR is an interactive process. With the current document-centered retrieval paradigm, interactive retrieval is treated as a sequence of independent simple retrieval decision making steps. However, it has been brought into attention that analysis of task-aware user sessions, which contain a sequence of requests submitted by a user to fulfill an information need, provides useful insight into the query behavior of the user. Third, most of present IR systems use keywords to query and index documents. However, this traditional keyword-based IR model provides little semantic context for the understanding of user information needs, which can lead to mismatch between query and document in search. Ideally, one would like to see the query and document match with each other, if they are topically relevant. Thus, the integration of semantic context according to the user's information need and the user's understanding of the documents into IR systems is needed to improve the IR performance. Fourth, organizations are now increasingly dealing with petabyte-scale collections of data. Context-sensitive and task-aware approaches become more challenging when dealing with big data. Hence, it is important to propose new algorithms and models that can effectively and efficiently process big data and implement the context-sensitive and task-aware approaches in big-data scenarios.****In a world where data are growing at extraordinary rates, there is a huge demand for searching and analyzing big data more accurately and effectively to discover useful information. This research program tackles the problem of searching and discovering useful information from big text data. The long-term objective of the proposed research is to overcome the limitations of the existing IR methods and formally develop a new retrieval paradigm called context-sensitive and task-aware information search for big data. In particular, (1) we will develop a new theoretical retrieval framework for capturing rich user information and providing personalized search results; (2) we will develop novel task-based retrieval methods for context-sensitive information retrieval to optimize the long-term retrieval utility over an entire retrieval session; (3) we will develop new models for automatically analyzing and searching big data to efficiently extract knowledge and perform semantic matching. **
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2020-07157
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$4.66万
-
财政年份:2022
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2020-07157
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$4.66万
-
财政年份:2021
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2020-07157
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$4.66万
-
财政年份:2020
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2015-03807
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.13万
-
财政年份:2019
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2015-03807
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.13万
-
财政年份:2017
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2015-03807
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.13万
-
财政年份:2016
-
负责人:Huang, Jimmy
-
依托单位:
Searching and Analyzing Big Data: Context-sensitive and Task-aware Approaches
-
批准号:RGPIN-2015-03807
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.13万
-
财政年份:2015
-
负责人:Huang, Jimmy
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: