Fundamental Limits of Data Analytics in Sociotechnical Systems

Fundamental Limits of Data Analytics in Sociotechnical Systems
复制标题

社会技术系统中数据分析的基本限制

DOI:
10.3389/fict.2016.00002
复制
发表时间:
2016
期刊:
Frontiers ICT
影响因子:
--
通讯作者:
L. Varshney
L. Varshney
中科院分区:
--
文献类型:
--
作者:
L. Varshney

文献摘要

被引文献

相似文献

在大数据时代,涉及人类和机器的信息系统正在部署在各种各样的社会环境中。许多人使用数据分析作为描述性、预测性和规范性任务的子组件,通常使用机器学习进行训练。然而,当分析组件被放置在大规模的社会技术系统中时,通常很难描述系统的运行情况,并用世界上相关的​​标准来衡量。在这里,我们提出了一种系统建模技术,将数据分析组件视为“嘈杂的黑匣子”或随机内核,它与基本随机分析一起提供了对基本性能限制的洞察。一个示例应用程序正在帮助优先考虑人们有限的注意力,其中学习算法使用噪声特征对任务进行排名,并且人们从排名列表中顺序进行选择。本文通过开发支持分析的顺序选择的随机模型来演示通用技术,使用顺序统计的伴随物得出基本限制,并评估系统范围性能指标(如筛选成本和所选对象的价值)的限制。还建立了与二分排名的样本复杂性的联系。
In the Big Data era, informational systems involving humans and machines are being deployed in multifarious societal settings. Many use data analytics as subcomponents for descriptive, predictive, and prescriptive tasks, often trained using machine learning. Yet when analytics components are placed in large-scale sociotechnical systems, it is often difficult to characterize how well the systems will act, measured with criteria relevant in the world. Here, we propose a system modeling technique that treats data analytics components as `noisy black boxes' or stochastic kernels, which together with elementary stochastic analysis provides insight into fundamental performance limits. An example application is helping prioritize people's limited attention, where learning algorithms rank tasks using noisy features and people sequentially select from the ranked list. This paper demonstrates the general technique by developing a stochastic model of analytics-enabled sequential selection, derives fundamental limits using concomitants of order statistics, and assesses limits in terms of system-wide performance metrics like screening cost and value of objects selected. Connections to sample complexity for bipartite ranking are also made.