Efficient support for multi-attribute top-k relational queries: a cost-based approach
Efficient support for multi-attribute top-k relational queries: a cost-based approach
批准号:
328087-2006
负责人:
Ayanso, Anteneh
金额:
$0.95万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2007
资助国家:
加拿大
项目状态:
已结题
起止时间:
2007-01-01 至 2008-12-31
中文摘要
“近似匹配”查询在文档和多媒体检索系统中非常常见。例如,在搜索引擎中,用户指定一组关键字,并期望返回与关键字相关的页面或文档的排名。相反,关系数据库管理系统(rdbms)中可用的查询方法被设计为只返回在指定选择条件下找到的结果。因此,在线商业应用程序(如产品推荐系统、价格比较和购物代理)的用户在搜索相关结果时经常面临指定属性值范围的挑战。通常,他们得到的结果要么太少,要么太多,而这些结果与他们的请求的相关性有限。这种传统的查询过程对用户来说非常令人沮丧,对系统来说效率极低。或者,上述应用程序的用户应该能够指定属性的目标值,并期望获得期望数量的结果的排序集,这些结果与所有属性中的指定值最匹配(例如,前10个最佳匹配)。在这种类型的查询(也称为top-k查询)中,结果不仅限于精确匹配,还包括感兴趣的目标值附近的接近匹配。本研究研究基于成本的策略,以便在rdbms中有效地支持这类查询。目标是提供在现有rdbms设计的技术限制范围内工作的方法,但避免对数据库进行完整的顺序扫描以获得top-k集。特别是,建议的研究引入了系统地结合相关性能成本因素及其潜在权衡的技术,以实现高效的top-k检索。该方法包括分析建模和广泛的计算和实验分析,使用广泛的实验设置的真实和合成数据集。
英文摘要
Querying for "approximate matches" is very common in document and multimedia retrieval systems. In search engines, for example, users specify a set of keywords and expect in return a ranking of relevant pages or documents related to the keywords. In contrast, the querying methods available in relational database management systems (RDBMSs) are designed to return only results found within specified selection conditions. Due to this, users of online business applications such as product recommendation systems, price comparison and shopping agents routinely face the challenge of specifying value ranges of attributes in search of relevant results. Often, they get either too few or too many results that are of limited relevance to their request. This conventional querying process is very frustrating for the user and extremely inefficient for the system. Alternatively, users of the above applications should be able to specify target values of attributes and expect to obtain a ranked set of a desired number of results that best match the specified values across all the attributes (e.g., the top 10 best matches). In this type of querying, also known as top-k querying, results are not limited to exact matches but include close matches around the target values of interest. This research studies cost-based strategies for efficient support of this class of queries in RDBMSs. The objective is to provide methods that work within the technical constraints of the existing design of RDBMSs but avoid a full sequential scan of the database to obtain the top-k set. In particular, the proposed research introduces techniques that systematically incorporate the relevant performance cost factors and their underlying trade-offs for efficient top-k retrieval. The methodology encompasses analytical modelling and extensive computational and experimental analyses using real and synthetic data sets over a wide range of experimental settings.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Efficient Strategies and Analytics Solutions for Social Media Targeting via Text Mining
-
批准号:522170-2018
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2018
-
负责人:Ayanso, Anteneh
-
依托单位:
Efficient support for multi-attribute top-k relational queries: a cost-based approach
-
批准号:328087-2006
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.95万
-
财政年份:2008
-
负责人:Ayanso, Anteneh
-
依托单位:
Efficient support for multi-attribute top-k relational queries: a cost-based approach
-
批准号:328087-2006
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.95万
-
财政年份:2006
-
负责人:Ayanso, Anteneh
-
依托单位:
国内基金
海外基金
两性离子载体(zwitterionic support)作为可溶性支载体在液相有机合成中的应用
-
批准号:21002080
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2010
-
负责人:霍聪德
-
依托单位:
微生物发酵过程的自组织建模与优化控制
-
批准号:60704036
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2007
-
负责人:高学金
-
依托单位:
基于Support Vector Machines(SVMs)算法的智能型期权定价模型的研究
-
批准号:70501008
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2005
-
负责人:曹丽娟
-
依托单位: