The Everlasting Database: Statistical Validity at a Fair Price

The Everlasting Database: Statistical Validity at a Fair Price
复制标题

DOI:
--
复制
发表时间:
2018-03
期刊:
--
影响因子:
--
通讯作者:
Blake E. Woodworth;V. Feldman;Saharon Rosset;N. Srebro
Blake E. Woodworth;V. Feldman;Saharon Rosset;N. Srebro
中科院分区:
其他
文献类型:
--
作者:
Blake E. Woodworth;V. Feldman;Saharon Rosset;N. Srebro

文献摘要

相似文献

在数据分析中处理自适应性的问题,无论是有意还是无意,都渗透到各个领域,包括ML挑战中的测试集过拟合和无效科学发现的积累。我们提出了一种机制,回答一个任意长的序列的潜在自适应统计查询,收费的价格为每个查询和使用的收益来收集额外的样本。至关重要的是,我们保证统计有效性,而不对查询如何生成进行任何假设。我们还以很高的概率确保$M$非自适应查询的成本是$O(\log M)$,而潜在自适应用户进行$M$查询的成本不依赖于任何其他查询是$O(\sqrt{M})$。
The problem of handling adaptivity in data analysis, intentional or not, permeates a variety of fields, including test-set overfitting in ML challenges and the accumulation of invalid scientific discoveries. We propose a mechanism for answering an arbitrarily long sequence of potentially adaptive statistical queries, by charging a price for each query and using the proceeds to collect additional samples. Crucially, we guarantee statistical validity without any assumptions on how the queries are generated. We also ensure with high probability that the cost for $M$ non-adaptive queries is $O(\log M)$, while the cost to a potentially adaptive user who makes $M$ queries that do not depend on any others is $O(\sqrt{M})$.