课题基金 / 基金详情

Predicting the Unlikely: Theory, Algorithms, and Applications

Predicting the Unlikely: Theory, Algorithms, and Applications
预测不可能发生的事情:理论、算法和应用
批准号:
0514973
负责人:
Alon Orlitsky
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-07-01 至 2009-06-30

项目摘要

项目成果

Alon Orlitsky的其他基金

相似基金

相关文献

中文摘要
翻译
摘要--------许多科学和工程工作要求基于观测数据样本估计概率和分布。当样本量相对于可能结果的数量很大时,例如,当一个有偏差的硬币被投掷多次时,估计就很简单了。然而,在许多应用中,与样本量相比,可能结果的数量很大。例如,在语言建模(用于压缩、语音识别和数据挖掘)中,单词和上下文的数量比手头的文本数量要大。在这种大字母体系下的估计要复杂得多,并且已经研究了两个多世纪。虽然已经导出了一些很好的估计器,例如以拉普拉斯、克里切夫斯基-特罗菲莫夫和古德-图灵命名的估计器,但为它们建立的最优性性质很少,而且在某些条件下,每个估计器的性能都很差。研究者们采用信息论的观点,对这些问题进行了系统的研究。它们集中在两个广泛的问题上,涉及到以下估计:(1)每个观察到的结果和尚未观察到的结果集合的概率;(2)基础分布,它不将概率与具体结果联系起来。对于每个问题,他们都寻求在实践中表现良好的估计算法,并且具有可证明的最优性属性,例如底层分布和估计分布之间的小Kullback-Leibler散度和其他度量。它们解决的问题既有理论的,例如在给定置信水平内估计底层分布所需的数据大小,也有计算的,关于派生算法的复杂性和顺序性。
英文摘要
Abstract--------Many scientific and engineering endeavors call for estimating probabilities and distributions based on an observed data sample. When the sample size is large relative to the number of possible outcomes, for example when a biased coin is tossed many times, estimation is simple. However, in many applications the number of possible outcomes is large compared to the sample size. For example, in language modeling - used in compression, speech recognition, and data mining - the number of words and contexts is large compared to the amount of text at hand. Estimation in this large-alphabet regime is much more complex, and has been studied for over two centuries. While some good estimators have been derived, for example those named afterLaplace, Krichevsky-Trofimov, and Good-Turing, very few optimality properties have been established for them, and each is known to perform poorly under some conditions. Adopting an information-theoretic viewpoint, the investigators undertake a systematic study of these issues. They concentrate on two broad problems, concerning the estimation of: (1) the probability of each observed outcome and of the collection of outcomes not yet observed; (2) the underlying distribution, which does not associate probabilities with specific outcomes. For each problem they seek estimation algorithms that perform well in practice and have provable optimality properties such as small Kullback-Leibler divergence and other metrics between the underlying and estimated distributions. The problems they address are both theoretical, for example the data size required to estimate the underlying distribution to within a given confidence level, and computational, regarding the complexity and sequentiality of the derived algorithms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CIF: Student Travel Support for the 2017 IEEE International Symposium on Information Theory
  • 批准号:
    1740960
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.0万
  • 财政年份:
    2017
  • 负责人:
    Alon Orlitsky
  • 依托单位:
CIF: SMALL: Information Theoretic Foundations of Data Science
  • 批准号:
    1619448
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2016
  • 负责人:
    Alon Orlitsky
  • 依托单位:
CIF: Medium: Collaborative Research: Learning in High Dimensions: From Theory to Data and Back
  • 批准号:
    1564355
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $59.85万
  • 财政年份:
    2016
  • 负责人:
    Alon Orlitsky
  • 依托单位:
Enhancing Education and Awareness of Shannon Theory
海外基金