课题基金 / 基金详情

Exploiting Special Structures in High-Dimensional Data Classification

Exploiting Special Structures in High-Dimensional Data Classification
在高维数据分类中利用特殊结构
批准号:
0505424
负责人:
Elizaveta Levina
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-06-01 至 2009-05-31

项目摘要

项目成果

Elizaveta Levina的其他基金

相似基金

相关文献

中文摘要
翻译
前言:提出的研究为高维数据分类开发了新的实用方法和算法以及理论结果。主要问题是,当测量变量的数量远远超过观测值的数量时,不可能准确地估计全协方差矩阵。这位研究人员此前已经证明,在这种情况下完全戒除依赖是一个更好的选择,但这意味着丢弃大量信息。这一问题将通过发展新的仅保留重要相关性信息的稀疏协方差估计来解决,并在判别分析的背景下从理论上研究它们的行为。在聚类和图分割方法的基础上,将开发新的实用的、计算高效的从数据中计算这些估计值的算法,并与现有的正则化技术进行比较。从理论上讲,渐近最优或接近最优的分类性能有望被证明。另一种降低问题规模和达到底层结构的方法是使用最近发展起来的非线性流形投影方法来降低数据维度,其目的是发现一种保留了数据中大部分信息的非线性低维嵌入。使用这些方法需要估计流形上的本征数据维度和邻域尺度,提出了两种新的严格估计,并分析了它们的统计性质。对这两个参数的仔细估计将改进目前在机器学习中使用的主要启发式方法,并增加多种投影方法对高维数据分类的适用性。该建议通过开发处理高维数据的新的理论和实践工具来应对现代世界中收集的大量数据带来的新挑战,特别是在一个观测的测量数量相对于观测数量较大的情况下。提出了一种新的稀疏估计器,它只包含与数据分类相关的信息。新的估计器还可以用于任何需要从有限的数据量中估计大协方差矩阵的问题,因此将对广泛的现代应用产生影响,例如基因表达数据的分类和分析、复杂的化学和物理实验的分析、遥感和医学成像等。
英文摘要
ABSTRACTProposed research develops new practical methodology and algorithms aswell as theoretical results for high-dimensional data classification. Themain issue is that when the number of measured variables exceeds by farthe number of observations, estimating the full covariance matrixaccurately is impossible. The investigator has previously shown thatignoring the dependence completely in such a situation is a better option,but it means discarding a lot of information. This problem will beresolved by developing new sparse covariance estimators which will onlyretain important dependence information, and studying their behaviortheoretically in the context of discriminant analysis. New practical,computationally efficient algorithms for computing these estimators fromdata will be developed on the basis of clustering and graph partitioningmethods and compared to existing regularization techniques. Theoretically,asymptotically optimal or near optimal classification performance isexpected to be demonstrated. Another way to reduce the size of theproblem and get to the underlying structure is to reduce the datadimension using recently developed nonlinear manifold projection methods,which aim to discover a nonlinear low-dimensional embedding preservingmost of the information contained in the data. Using these methodsrequires estimating intrinsic data dimension and neighborhood scale on amanifold, and new rigorous estimators for both are proposed, along with ananalysis of their statistical properties. Careful estimation of these twoparameters will improve on the current mostly heuristic methods used inmachine learning and increase applicability of manifold projection methodsfor high-dimensional data classification.This proposal addresses the new challenges posed by the massive amounts ofdata collected in the modern world by developing new theoretical andpractical tools for dealing with high-dimensional data, particularly withthe situation when the number of measurements taken for one observation islarge relative to the number of observations. New sparse estimators ofdependence structure in such data are developed, which only contain theinformation relevant for data classification. The new estimators can alsobe used in any problem where large covariance matrices need to beestimated from limited amount of data, and hence will have an impact on awide range of modern applications, such as classification and analysis ofgene expression data, analysis of complex chemical and physicalexperiments, remote sensing, and medical imaging, among others.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FRG: Collaborative Research: Flexible Network Inference
Multivariate Analysis for Samples of Networks
RTG: Understanding dynamic big data with complex structure
Conference proposal: From Industrial Statistics to Data Science
国内基金
海外基金
非阶化Hamiltonial型和Special型李代数的表示
  • 批准号:
    10701002
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    15.0万元
  • 批准年份:
    2007
  • 负责人:
    赵玉凤
  • 依托单位: