Files of a Feather Flock Together? Measuring and Modeling How Users Perceive File Similarity in Cloud Storage

Files of a Feather Flock Together? Measuring and Modeling How Users Perceive File Similarity in Cloud Storage
复制标题

DOI:
10.1145/3404835.3462845
复制
发表时间:
2021-07
期刊:
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Will Brackenbury;Galen Harrison;K. Chard;Aaron J. Elmore;Blase Ur
Will Brackenbury;Galen Harrison;K. Chard;Aaron J. Elmore;Blase Ur
中科院分区:
其他
文献类型:
--
作者:
Will Brackenbury;Galen Harrison;K. Chard;Aaron J. Elmore;Blase Ur

文献摘要

相似文献

先前的研究表明,用户通过相似性的视角来概念化个人数字文件收藏的组织。然而,尚不清楚在实际文件集合中相似文件实际上在多大程度上彼此靠近(例如,在同一目录中),或者利用文件相似性是否可以改进无组织文件集合的信息检索和组织。为此,我们进行了一项在线研究,将对 50 个 Google Drive 和 Dropbox 用户的云帐户进行自动分析,并进行一项询问这些帐户中的文件对的调查。我们发现,位于文件层次结构不同部分的许多文件在参与者的感知方式以及算法可提取的特征方面都很相似。参与者通常希望共同管理类似的文件(例如,删除一个文件意味着删除另一个文件),即使它们在文件层次结构中相距甚远。为了进一步理解这种关系,我们建立了回归模型,找到了几个可通过算法提取的文件特征来预测人类对文件相似性和所需文件共同管理的看法。我们的研究结果为利用文件相似性根据用户之前与类似文件的交互自动推荐访问、移动或删除操作铺平了道路。
Prior work suggests that users conceptualize the organization of personal collections of digital files through the lens of similarity. However, it is unclear to what degree similar files are actually located near one another (e.g., in the same directory) in actual file collections, or whether leveraging file similarity can improve information retrieval and organization for disorganized collections of files. To this end, we conducted an online study combining automated analysis of 50 Google Drive and Dropbox users' cloud accounts with a survey asking about pairs of files from those accounts. We found that many files located in different parts of file hierarchies were similar in how they were perceived by participants, as well as in their algorithmically extractable features. Participants often wished to co-manage similar files (e.g., deleting one file implied deleting the other file) even if they were far apart in the file hierarchy. To further understand this relationship, we built regression models, finding several algorithmically extractable file features to be predictive of human perceptions of file similarity and desired file co-management. Our findings pave the way for leveraging file similarity to automatically recommend access, move, or delete operations based on users' prior interactions with similar files.