Expressive retrieval from XML documents

Expressive retrieval from XML documents
复制标题

DOI:
10.1145/383952.383982
复制
发表时间:
2001-09
期刊:
--
影响因子:
--
通讯作者:
Taurai Tapiwa Chinenyanga;N. Kushmerick
Taurai Tapiwa Chinenyanga;N. Kushmerick
中科院分区:
其他
文献类型:
--
作者:
Taurai Tapiwa Chinenyanga;N. Kushmerick

文献摘要

被引文献

相似文献

XML 作为结构化文档/数据的标准交换格式的出现引发了许多 XML 查询语言提案。然而,其中一些语言不支持基于文本相似性的信息检索式排名查询。这些查询语言已经有多种扩展来支持关键字搜索,但是生成的查询语言无法表达诸如“查找具有相似标题的书籍和 CD”之类的查询。这些扩展要么使用关键字作为纯粹的布尔过滤器,要么只能计算数据值和常量之间的相似度,而不是两个数据值之间的相似度。我们提出了 ELIXIR,一种用于 \textbf{\underline{X}}ML \textbf{\underline{i}}nformation \textbf{\underline{r}} 检索的 \textbf{\underline{e}}xpressive 和 \textbf{\underline{e}}fficient\textbf{\underline{l}} 语言,扩展了查询语言 XML-QL \cite{deutsch-www8,deutsch-deb99} 带有文本相似性运算符。 ELIXIR 是一种通用 XML 信息检索语言,具有足够的表达能力来处理上述查询。我们用于回答 ELIXIR 查询的算法将原始 ELIXIR 查询重写为一系列生成中间关系数据的 XML-QL 查询,并使用关系数据库技术有效地评估该中间数据的相似性运算符,生成一个 XML 文档,其中节点按相似性排名。我们的实验表明,我们的原型可以很好地适应 XML 数据的大小和查询的复杂性。
The emergence of XML as a standard interchange format for structured documents/data has given rise to many XML query language proposals. However, some of these languages do not support information retrieval-style ranked queries based on textual similarity. There have been several extensions to these query languages to support keyword search, but the resulting query languages cannot express queries such as``find books and CDs with similar titles''. Either these extensions use keywords as mere boolean filters, or similarities can be calculated only between data values and constants rather than two data values. We propose ELIXIR, an \textbf{\underline{e}}xpressive and \textbf{\underline{e}}fficient\textbf{\underline{l}}anguage for \textbf{\underline{X}}ML \textbf{\underline{i}}nformation \textbf{\underline{r}}etrieval that extends the query language XML-QL \cite{deutsch-www8,deutsch-deb99} with a textual similarity operator. ELIXIR is a general-purpose XML information retrieval language, sufficiently expressive to handle the above query. Our algorithm for answering ELIXIR queries rewrites the original ELIXIR query into a series of XML-QL queries that generate intermediate relational data, and uses relational database techniques to efficiently evaluate the similarity operators on this intermediate data, yielding an XML document with nodes ranked by similarity. Our experiments demonstrate that our prototype scales well with the size of the XML data and complexity of the query.