课题基金 / 基金详情

Statistical Methods for Prediction of Individual Sequences

Statistical Methods for Prediction of Individual Sequences
预测个体序列的统计方法
批准号:
0707060
负责人:
Peter Bartlett
金额:
$23.72万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-07-01 至 2011-06-30

项目摘要

项目成果

Peter Bartlett的其他基金

相似基金

相关文献

中文摘要
翻译
在许多出现的预测问题中,例如,在计算机安全和计算金融中,生成数据的过程最好被建模为与预测者竞争的对手。该研究项目的主要目标是分析和设计对抗环境中复杂预测问题的统计方法:预测策略必须几乎与某些比较类中的最佳策略一样准确地预测任何单个序列。研究的重点是基于数据概率模型的统计方法,因为(a)这些方法在实践中经常使用;(b)有证据表明这些方法在对抗环境中表现良好;(c)利用为概率设置开发的计算效率高的方法有很好的前景,这对高维问题特别重要;(d)对抗性设置中的积极结果应该更好地理解统计方法在概率设置中的稳健性。研究目标是:开发分析统计方法性能的技术,如贝叶斯方法,用于预测单个序列;提高对测量这些方法所使用的概率模型复杂性的适当方法的理解;阐明计算简化(如经验贝叶斯方法和MAP估计)对这些预测策略性能的影响;因此,为复杂的对抗性预测问题开发计算有效的统计方法的设计方法。对于许多估计和预测问题,将生成数据的过程建模为与预测器竞争的对手是合适的。这样的问题在信息技术领域很常见。例如,在垃圾邮件过滤问题中,目标是将传入的电子邮件标记为合法或垃圾邮件,但与此同时,垃圾邮件发送者试图设计电子邮件消息以躲过这些垃圾邮件过滤器。因此,预测问题是一个两人重复博弈。类似的决策问题出现在计算机网络安全(例如,决定网络流量是正常的还是拒绝服务攻击的结果的问题)和互联网搜索(例如,决定一个高链接的网页是否真正权威,应该有高页面排名)。在这些问题中,预测策略所看到的部分数据被对手选择,目的是欺骗预测策略。这些对抗性问题在金融应用中也很常见。假设从金融时间序列分析中得出的预测用于优化投资组合中的资本配置。然后,在短期内,为了其他市场参与者的利益,采取行动减少投资组合的回报,从而使预测变得不准确。统计方法(如贝叶斯方法)通常应用于这些对抗性预测问题,尽管它们是为数据的概率模型而不是对抗性模型而设计的。本研究项目旨在开发技术,以了解这些预测方法性能的固有局限性,从而为设计更强大的预测方法提供信息。重点研究了适用于实际中出现的复杂、高维预测问题的预测精度和计算效率。
英文摘要
In many prediction problems that arise, for example, in computer security and computational finance, the process generating the data is best modeled as an adversary with whom the predictor competes. The broad goal of this research project is the analysis and design of statistical methods for complex prediction problems in an adversarial setting: a prediction strategy must predict any individual sequence almost as accurately as the best strategy in some comparison class. The research focuses on statistical methods, which are based on probabilistic models of the data, since (a) these methods are commonly used in practice; (b) there is evidence that these methods perform well in adversarial settings; (c) there are good prospects of exploiting the computationally efficient approaches that have been developed for probabilistic settings, which is especially important for high-dimensional problems; and (d) positive results in adversarial settings should provide better understanding of the robustness of statistical methods in probabilistic settings. The research aims are: to develop techniques for analyzing the performance of statistical methods, such as Bayesian methods, for the prediction of individual sequences; to improve the understanding of appropriate ways to measure the complexity of a probability model used by these methods; to elucidate the impact of computational simplifications (such as empirical Bayes approaches and MAP estimates) on the performance of these prediction strategies; and hence to develop design methodologies for computationally efficient statistical methods for complex adversarial prediction problems.There are many estimation and prediction problems for which it is appropriate to model the process generating the data as an adversary with whom the predictor competes. Such problems are common in information technology. For instance, in the problem of spam filtering, the aim is to label incoming email as either legitimate or spam, but at the same time, spammers try to design email messages that slip past these spam filters. Thus, the prediction problem is a two-player repeated game. Similar decision problems arise in computer network security (for instance, the problem of deciding whether network traffic is normal or the result of a denial-of-service attack) and in internet search (for instance, deciding if a highly linked web page is genuinely authoritative and should have a high page rank). In these problems, some fraction of the data seen by the prediction strategy is chosen by an adversary who aims to fool the prediction strategy. These adversarial problems are also common in financial applications. Suppose that the predictions that emerge from a financial time series analysis are used to optimize the allocation of capital across a portfolio. Then, in the short term, it is in the interests of other market players to act so as to diminish the returns of the portfolio, and thus render the predictions inaccurate. It is common for statistical methods, such as Bayesian methods, to be applied to these adversarial prediction problems, despite the fact that they are designed for probabilistic, rather than adversarial, models of the data. This research project aims to develop techniques to understand the inherent limitations on the performance of these prediction methods, and hence inform the design of more powerful prediction methods. It focuses on the predictive accuracy and computational efficiency of methods that are suitable for the complex and high-dimensional prediction problems that arise in practise.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Conference: Women-in-Theory Workshop
  • 批准号:
    2227705
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2022
  • 负责人:
    Peter Bartlett
  • 依托单位:
Collaboration on the Theoretical Foundations of Deep Learning
  • 批准号:
    2031883
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $500.0万
  • 财政年份:
    2020
  • 负责人:
    Peter Bartlett
  • 依托单位:
Foundations of Data Science Institute
  • 批准号:
    2023505
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $590.03万
  • 财政年份:
    2020
  • 负责人:
    Peter Bartlett
  • 依托单位:
RI: AF: Small: Optimizing probabilities for learning: sampling meets optimization
  • 批准号:
    1909365
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2019
  • 负责人:
    Peter Bartlett
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data