20.453j / 2.771j / Hst.958j Biomedical Information Technology Relational Databases for Querying Xml Documents: Limitations and Opportunities

20.453j / 2.771j / Hst.958j Biomedical Information Technology Relational Databases for Querying Xml Documents: Limitations and Opportunities
复制标题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
J. Shanmugasundaram;Kristin Tufte;G. He;Chun Zhang;D. DeWitt;J. Naughton
J. Shanmugasundaram;Kristin Tufte;G. He;Chun Zhang;D. DeWitt;J. Naughton
中科院分区:
其他
文献类型:
--
作者:
J. Shanmugasundaram;Kristin Tufte;G. He;Chun Zhang;D. DeWitt;J. Naughton

文献摘要

被引文献

相似文献

XML正迅速成为万维网中表示数据的主要标准。允许用户有效地利用存储在XML文档中的数据的复杂查询引擎对于充分利用XML的能力至关重要。虽然最近有大量的活动提出了新的半结构化数据模型和查询语言,本文探讨了更保守的方法,使用传统的关系数据库引擎处理符合文档类型描述符(DTD)的XML文档。为此,我们已经开发了算法,并实现了一个原型系统,将XML文档转换为关系元组,将XML文档的半结构化查询转换为表上的SQL查询,并将结果转换为XML。我们已经定性地评估了这种方法,使用几个真实的DTD从不同的领域。事实证明,关系方法可以处理XML数据上的半结构化查询的大部分(但不是全部)语义,但可能仅在某些情况下有效。我们确定了这些限制的原因,并提出了某些扩展的关系允许免费复制本材料的全部或部分被授予,前提是复制品不是为了直接的商业利益而制作或分发的,VLDB版权声明和出版物的标题及其日期出现,并通知复制是由超大型数据库基金会许可的。复制或重新发布需要付费和/或获得Endowment模型的特殊许可,这将使其更适合于处理XML文档上的查询。
XML is fast emerging as the dominant standard for representing data in the World Wide Web. Sophisticated query engines that allow users to effectively tap the data stored in XML documents will be crucial to exploiting the full power of XML. While there has been a great deal of activity recently proposing new semi-structured data models and query languages for this purpose, this paper explores the more conservative approach of using traditional relational database engines for processing XML documents conforming to Document Type Descriptors (DTDs). To this end, we have developed algorithms and implemented a prototype system that converts XML documents to relational tuples, translates semi-structured queries over XML documents to SQL queries over tables, and converts the results to XML. We have qualitatively evaluated this approach using several real DTDs drawn from diverse domains. It turns out that the relational approach can handle most (but not all) of the semantics of semi-structured queries over XML data, but is likely to be effective only in some cases. We identify the causes for these limitations and propose certain extensions to the relational Permission to copy without fee all or part of this material is granted provided that the copies are not made or distributed for direct commercial advantage, the VLDB copyright notice and the title of the publication and its date appear, and notice is given that copying is by permission of the Very Large Data Base Endowment. To copy otherwise, or to republish, requires a fee and/or special permission from the Endowment model that would make it more appropriate for processing queries over XML documents.