Analysis of Long Term File Reference Patterns for Application to File Migration Algorithms

Analysis of Long Term File Reference Patterns for Application to File Migration Algorithms
复制标题

应用于文件迁移算法的长期文件参考模式分析

DOI:
10.1109/tse.1981.230843
复制
发表时间:
1981
影响因子:
7.4
通讯作者:
A. Smith
A. Smith
中科院分区:
计算机科学1区
文献类型:
--
作者:
A. Smith

文献摘要

被引文献

相似文献

在大多数大型计算机安装中,文件由系统自动和/或在用户的指示下在在线磁盘和大容量存储(磁带、集成大容量存储设备)之间移动。本文提出并分析了长期文件参考数据,为文件迁移算法的构建奠定了基础。具体地说,我们通过分析13个月的文件参考数据,检查了斯坦福直线加速器中心(SLAC)计算机安装的在线用户(主要是文本编辑)数据集的使用情况。我们发现大多数文件被使用的次数很少。在那些被足够频繁地使用以至于可以检查其参考模式的参数中,我们发现:1)大约三分之一的参考速率在其一生中显示出下降的参考速率,2)其余的参考间隔中,极少数(约5%)显示出相关的参考间隔,以及3)参考间隔(以天为单位)似乎比伯努利过程中出现的更偏斜。因此,在所有充分活跃的文件中,约有三分之二的文件似乎被引用为具有扭曲的相互引用分布的续订过程。大量其他文件引用统计信息(文件生存期、干扰分布、矩、平均值、使用次数/文件、文件大小、文件引用速率等)进行了计算并给出了计算结果。从头到尾,统计测试都会被描述和解释。我们对文件引用模式的分析结果被应用在一篇配对论文中,用于开发和比较评估文件迁移算法。
In most large computer installations files are moved between on-line disk and mass storage (tape, integrated mass storage device) either automatically by the system and/or at the direction of the user. In this paper we present and analyze long term file reference data in order to develop a basis for the construction of algorithms for file migration. Specifically, we examine the use of the on-line user (primarily text editor) data sets at the Stanford Linear Accelerator Center (SLAC) computer installation through the analysis of 13 months of file reference data. We find that most files are used very few times. Of those that are used sufficiently frequently that their reference patterns may be examined, we find that: 1) about a third show declining rates of reference during their lifetime, 2) of the remainder, very few (about 5 percent) show correlated interreference intervals, and 3) interreference intervals (in days) appear to be more skewed than would occur with the Bernoulli process. Thus, about two-thirds of all suffi1ciently active files appear to be referenced as a renewal process with a skewed interreference distribution. A large number of other file reference statistics (file lifetimes, interference distributions, moments, means, number of uses/ file, file sizes, file rates of reference, etc.) are computed and presented. Throughout, statistical tests are described and explained. The results of our analysis of file reference patterns are applied in a companion paper to the development and comparative evaluation of file migration algorithms.