The case for efficient file access pattern modeling

The case for efficient file access pattern modeling
复制标题

DOI:
10.1109/hotos.1999.798371
复制
发表时间:
1999-03
期刊:
Proceedings of the Seventh Workshop on Hot Topics in Operating Systems
影响因子:
--
通讯作者:
Thomas M. Kroeger;D. Long
Thomas M. Kroeger;D. Long
中科院分区:
其他
文献类型:
--
作者:
Thomas M. Kroeger;D. Long

文献摘要

被引文献

相似文献

大多数现代I/O系统独立地对待每个文件访问。然而,计算机系统中的事件是由程序驱动的。因此,对文件的访问以一致的模式发生,并且绝不是独立的。其结果是现代I/O系统忽略了有用的信息。通过对文件系统活动的跟踪,我们发现文件访问与之前的访问密切相关。事实上,一个简单的最后-后继模型(预测每次文件访问之后将会是最后一次访问的同一个文件)成功预测下一个文件的概率为72%。我们将先前提出的两种文件访问预测模型与此基线模型进行比较,并在准确性和状态空间的高开销方面看到了鲜明的对比。然后我们增强其中一个模型来解决模型空间需求的问题。这个新模型能够在最后一个后继模型的恒定因子(相对于文件数量)的状态空间内工作时,将最后一个后继模型的准确性提高10%。虽然这项工作的动机是使用文件关系进行l/O预取,但有关文件访问模式可能性的信息还有其他几种用途,例如磁盘布局和用于断开连接操作的文件集群。
Most modern I/O systems treat each file access independently. However events in a computer system are driven by programs. Thus, accesses to files occur in consistent patterns and are by no means independent. The result is that modern I/O systems ignore useful information. Using traces of file system activity we show that file accesses are strongly correlated with preceding accesses. In fact, a simple last-successor model (one that predicts each file access will be followed by the same file that followed the last time it was accessed) successfully predicted the next file 72% of the time. We examine the ability of two previously proposed models for file access prediction in comparison to this baseline model and see a stark contrast in accuracy and high overheads in state space. We then enhance one of these models to address the issues of model space requirements. This new model is able to improve an additional 10% on the accuracy of the last-successor model, while working within a state space that is within a constant factor (relative to the number of files) of the last successor model. While this work was motivated by the use of file relationships for l/O prefetching, information regarding the likelihood of file access patterns has several other uses such as disk layout and file clustering for disconnected operation.