The design of a relational database for large-scale bibliographic retrieval

The design of a relational database for large-scale bibliographic retrieval
复制标题

大规模书目检索关系数据库的设计

DOI:
--
复制
发表时间:
1996
影响因子:
1.8
通讯作者:
R. Green
R. Green
中科院分区:
管理学4区
文献类型:
--
作者:
R. Green

文献摘要

被引文献

相似文献

一个完全规范化的关系型书目数据库,承诺救济的更新,插入和删除异常,困扰书目数据库使用(美国)MARC格式内部。提出了一个基于实体关系模型的全规模书目数据库(包括书目数据、权威数据、馆藏数据和分类数据)的概念设计。这种设计很容易转换为逻辑关系设计。讨论了格式整合的处理以及知识和书目描述层次之间以及集体和个人描述层次之间的区别。不幸的是,书目数据的复杂性导致了关系方法的语义完整性与规范化和分解的低效率之间的紧张关系。折衷办法的困境概述。根据日期,“不可否认的是,关系的方法代表了当今市场的主导趋势,'关系模型'。. .是数据库领域整个历史上最重要的发展”(1)尽管这种数据模型在更大的数据库世界中很流行,但大型书目数据库通常保留基于复杂但单一的MARC记录的数据结构。虽然某种程度的关系是隐含的-例如,在使用USMARC书目,权威,馆藏和分类格式,只有有限的直接记录间的联系发生;(2)书目数据库世界仍然是相当多的记录导向比关系导向。(本文中提到的MARC通常是USMARC专用的。然而,这里提出的更大的结论也适用于其他MARC实现。)鉴于关系数据库的普遍有效性,根据关系数据库原则重新设计书目数据库的想法在文献中不时出现也就不足为奇了。(3)1994年秋季,马里兰州大学图书馆和信息服务学院(CLIS)举行了一次研讨会,从类似的前提出发,着手采用实体-关系(ER)模型建立一个大型书目数据库的基本逻辑设计,以期最终将基于ER的概念方案转换为关系数据库。这一承诺的结果在本文中报告和讨论。背景MARC和数据库设计在书目领域,MARC家族构成了无与伦比的标准。机读目录格式是为机读形式的书目数据的传输而设计的,特别是在磁带上,它首先是一种通信格式。由于缺乏合适的替代品,它们也被用作存储格式。在标准的三级数据库体系结构中,MARC格式通常与外部模式最为接近,它区分了内部模式(物理数据存储结构)、概念模式(数据的逻辑、社区范围的视图)和外部模式(数据的用户视图,特别是在输出报告或屏幕显示方面)。这种三级数据库体系结构支持数据独立性的理想,即在数据库的一个级别或模式中进行更改而无需复制数据库其他级别中的更改的能力。(6)一方面,在输入和/或输出端使用MARC记录作为通信格式既不规定书目数据库的逻辑视图,也不规定其内部存储结构。另一方面,内部存储结构必须与MARC兼容,这样来自MARC记录的数据可以转换为与数据库中已有的数据一致,数据库中的数据可以转换为MARC记录以供输出; Llorens和Trenor介绍了这样一种系统。…
A fully normalized relational bibliographic database promises relief from the update, insertion, and deletion anomalies that plague bibliographic databases using (US) MARC formats internally. The conceptual design of a full-scale bibliographic database (including bibliographic, authority, holdings, and classification data) is presented, based on entity-relationship modeling. This design translates easily into a logical relational design. The treatment of format integration and the differentiation between the intellectual and bibliographic levels of description and between collective and individual levels of description are discussed. Unfortunately, the complexities of bibliographic data result in a tension between the semantic integrity of the relational approach and the inefficiencies of normalization and decomposition. Compromise approaches to the dilemma are outlined. According to Date, "It is undeniable that the relational approach represents the dominant trend in the marketplace today, and that the `relational model' . . . is the single most important development in the entire history of the database field"(1) Despite the popularity of this data model in the larger database world, large-scale bibliographic databases have typically retained data structures based on the complex, but unitary MARC record. While some degree of relationality is implied--for example, in the use of the USMARC bibliographic, authority, holdings, and classification formats-only limited direct interrecord linkage occurs;(2) the bibliographic database world is still considerably more record-oriented than relation-oriented. (Reference to MARC in this article is often USMARC-specific. Nevertheless, the larger conclusions presented here should hold also for other MARC implementations.) Given the general effectiveness of relational databases, it is not surprising that the idea of redesigning bibliographic databases according to relational database principles has cropped up from time to time in the literature.(3) Starting from a similar premise, a seminar conducted in the College of Library and Information Services (CLIS) at the University of Maryland in fall 1994(4) undertook to establish the basic logical design of a large-scale bibliographic database using the entity-relationship (ER) model, with a view to the eventual conversion of the ER-based conceptual schemes into a relational database. The results of that undertaking are reported and discussed in this article. Background MARC and Database Design In the bibliographic world, the MARC family constitutes an unrivaled standard. Designed for the transfer of bibliographic data in machine-readable form, specifically on magnetic tape, the MARC formats are, first and foremost, communication formats. For lack of suitable alternatives, they have also been used as storage formats. In the standard three-level database architecture,(5) which distinguishes among internal schemes (physical data storage structures), conceptual schemes (logical, community-wide views of the data), and external schemes (user views of the data, especially in terms of output reports or screen displays), the MARC formats generally correspond most closely to external schemes. This three-level database architecture supports the ideal of data independence, the capacity to make changes in one level or schema of the database without having to replicate changes in other levels of the database.(6) On the one hand, the use of MARC records as a communications format on the input and / or output side dictates neither the logical view of a bibliographic database nor its internal storage structure. On the other hand, the internal storage structure must be MARC-compatible, so that data coming in from a MARC record can be transformed to be consistent with data already in the database and data in the database can be transformed into a MARC record for output; Llorens and Trenor introduce just such a system. …