Personalised Federated Search of the Deep Web (NEMO)
Personalised Federated Search of the Deep Web (NEMO)
批准号:
EP/F060475/1
负责人:
Fabio Crestani
金额:
$24.06万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --
中文摘要
分布式信息检索(DIR),也称为基于内容的联合搜索,致力于使用户能够使用表达其信息需求的自然语言查询,根据其语义内容来查找非结构化或结构不良的文档。与标准信息检索的不同之处在于,文档包含在多个不同的分布式资源中,每个资源都有不同的检索引擎,当有如此多的资源可用时,用户面临的第一个信息访问任务是资源选择。这是一项无效的人工任务,因为用户往往不知道每个资源在数量、质量、信息类型、来源和可能的相关性方面的内容。人们需要准确的自动资源选择工具来帮助他们完成这项任务,但资源选择需要准确的资源描述。这些描述可以手动或自动建立,在非合作资源的情况下是一个真正的问题,即在资源通过查询其搜索引擎来访问其内容的情况下,但不提供关于其档案内容的任何信息。这是深度或隐藏网络中的资源的典型情况。粗略估计,这些资源的规模是可见网络的400倍。一旦选择了资源并将查询转发给它们,每个资源返回的结果必须通过一个称为结果融合的过程进行合并,从而生成一个单一的结果排序列表并呈现给用户,试图最大化总体检索质量。过去10年的大量研究表明,通过使系统适应特定的用户任务和需求,可以极大地提高IR系统的效率。这使得能够在考虑到放置用户需求的上下文的情况下满足用户信息需求,从而使与系统的交互个性化。虽然已经有大量关于个性化和上下文相关的信息检索的工作,但DIR研究领域尚未考虑个性化问题。我们认为,自适应DIR的设计方法使DIR系统能够自动适应用户的需求和任务,与标准IR的设计方法一样重要,并将带来类似的好处。该项目致力于设计、实现和测试个性化的基于内容的联合搜索模型,这些模型将应用于从Deep Web检索信息。该项目的目标将通过设计、实现和测试高级资源描述、资源选择和结果融合方法来实现,这些方法可以根据用户任务和用户需求自动个性化,并且专门设计用于访问Deep Web中的信息(即,在非合作和不同的资源中)。据我们所知,这项建议是第一次尝试处理这一非常重要和即将到来的研究领域,并将推动直接投资研究的最新水平以及直接投资研究进程的每一个组成部分。这项工作的结果还将对系统的设计产生相当大的商业影响,这些系统将使人们能够访问目前大多尚未开发的深网的大量资源。
英文摘要
Distributed Information Retrieval (DIR), also known as content-based federated search, is concerned with enabling a user to find unstructured or poorly structured documents by their semantic content using natural language queries expressing his/her information needs. The difference with standard Information Retrieval (IR) is that documents are contained in a number of heterogeneous distributed resources, each with its own different retrieval engine.When so many resources are available, the first information access task the user faces is resource selection. This is an ineffective manual task as users are often unaware of the contents of each resource in terms of quantity, quality, information type, provenance and likely relevance. People need accurate automatic resource selection tools to assist them in this task, but resource selection requires accurate resource descriptions. These descriptions can be built either manually or automatically, and are a real problem to derive in the case of non-cooperative resources, that is in the case of resources that enable access to their content by querying their search engines, but that do not provide any information about the content of their archives. This is the typical case for resources in the Deep or Hidden Web. A rough estimate put the size of these resources at 400 times that of the Visible Web. Once the resources have been selected and the query forwarded to them, the results returned by each one of them have to be merged by a process called results fusion, so that a single ranked list of results is produced and presented to the user, trying to maximise the overall retrieval quality.A large body of research in the last 10 years has shown that the effectiveness of IR systems can be greatly improved by adapting the system to the specific user tasks and needs. This enables to satisfy the user information need taking into consideration the context in which the user need is placed, so personalising the interaction with the system. While a large body of work already exists for personalised and context dependent IR, the DIR research area has not yet considered issues of personalisation. We believe that designing methods for adaptive DIR, that enable a DIR system to automatically adapt to the user needs and task, is as important as designing them for standard IR and will bring similar benefits. This project is concerned with designing, implementing and testing models of personalised content-based federated search that will be applied to retrieving information from the Deep Web.The objective of this project will be achieved by designing, implementing and testing advanced resource description, resource selection and results fusions methods that can be automatically personalised to the user task and user needs and that are specifically designed to access information held in the Deep Web (i.e.~in non-ccoperative and heterogeneous resources). To thebest of our knowledge this proposal is the first to attempt to tackle this very important and up-coming area of research, and will advance the state of the art of DIR and of each and every component of the DIR process. The results of this work will also have considerable commercial implications for the design of systems that will enable access to the vast wealth of resources of the Deep Web that are, currently, mostly untapped.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Metadata harvesting for content-based distributed information retrieval
用于基于内容的分布式信息检索的元数据收集
DOI:
10.1002/asi.20694
发表时间:
2007
期刊:
Journal of the American Society for Information Science and Technology
影响因子:
--
作者:
[Simeoni F]
通讯作者:
Simeoni F
Cytogenetic identification and molecular marker development for the novel stripe rust-resistant wheat-Thinopyrum intermedium translocation line WTT11.
新型抗条锈病小麦-Thinopyrum intermedium易位系WTT11的细胞遗传学鉴定和分子标记开发。
DOI:
10.1007/978-3-642-15464-5_60
发表时间:
2021
期刊:
aBIOTECH
影响因子:
--
作者:
[Yang G]
通讯作者:
Yang G
Measuring the likelihood property of scoring functions in general retrieval models
测量一般检索模型中评分函数的似然属性
DOI:
10.1002/asi.21048
发表时间:
2009
期刊:
Journal of the American Society for Information Science and Technology
影响因子:
--
作者:
[Bache R]
通讯作者:
Bache R
SPIRE'06 Symposium in Glasgow: Support for Student Attendance
-
批准号:EP/D078598/1
-
项目类别:Research Grant
-
资助金额:$2.17万
-
财政年份:2006
-
负责人:Fabio Crestani
-
依托单位:
An Interactive Modus Operandi Visualisation System Integrating Geographical Information For Suspect Prioritisation and Investigation Management(iMOV)
-
批准号:EP/D040639/1
-
项目类别:Research Grant
-
资助金额:$12.54万
-
财政年份:2006
-
负责人:Fabio Crestani
-
依托单位:
海外基金