课题基金 / 基金详情

Significance Based Procedures for Mining and Prediction of Large Data Sets

Significance Based Procedures for Mining and Prediction of Large Data Sets
基于显着性的大数据集挖掘和预测程序
批准号:
0907177
负责人:
Andrew Nobel
金额:
$21.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2013-08-31

项目摘要

项目成果

Andrew Nobel的其他基金

相似基金

相关文献

中文摘要
翻译
探索性方法在理解大型数据集方面发挥着关键作用,无论其来源如何,并且通常是分析数据集的第一步。研究者正在研究在高维数据中识别模式或规律的探索性数据挖掘方法的开发和使用。他的研究重点是识别可能由多种测量技术产生的大型数据集中的样本变量关联问题。在实验数据以矩形矩阵形式表示的典型情况下,样本-变量关联对应于数据矩阵的不同子矩阵。研究者正在开发一种统计原则,基于意义的方法来寻找数据矩阵的大平均子矩阵的问题,使用简单的迭代算法。该算法适用于实值和分类数据矩阵。除了基本方法之外,研究者正在开发一些扩展,包括数据驱动的零模型,该模型包含变量之间的依赖性,同时应用多种测量技术产生的数据,以及将基本方法应用于预测问题,如分类,回归和生存分析。此外,研究者正在开发基础理论来支持算法的使用,并评估不同null模型下数据矩阵的结构。这些方法的开发和应用是在几个生物医学研究小组的密切合作下进行的。特别是,新的数据挖掘方法正被纳入合作科学家用于识别和评估正在进行的涉及乳腺癌、脑癌和肺癌的实验中重要的样本变量关联的软件中。现在,在许多科学实验领域,特别是在癌症等人类疾病的基因水平研究中,大型数据集很常见。在这样的研究中,经常会遇到包含数百到数千个样本的实验,以及对每个样本进行数万到数百万次测量的实验。大型数据集是从传统的假设驱动型科学研究转向数据驱动型研究的趋势的一部分,在这种趋势中,研究人员探索大型数据集的模式或规律,与主题专业知识相结合,产生可以通过更传统的方法进行检验的假设。研究者正在研究一种探索性方法,该方法可以识别大数据集中样本和变量之间的统计显著关联,这种关联可以产生可检验的科学假设。研究者正在开发的方法计算效率高,并基于既定的统计原则,特别是统计显著性的概念。研究者还在研究如何将基本探索性方法应用于多种测量技术产生的数据,以及如何将基本探索性方法应用于分类、生存分析等统计问题。这些活动是作为合作研究计划的一部分进行的,该计划涉及来自统计、生物和医学科学的教师和学生的持续互动。研究者开发的探索方法被整合到合作科学家的基本探索工具中,并且是分析几个新的,以前未分析的数据集的一个组成部分。
英文摘要
Exploratory methods play a critical role in the understanding of large data sets, regardless of their origin, and are typically the first step in their analysis. The investigator is studying the development and use of exploratory, data-mining methods that identify patterns or regularities in high-dimensional data. The specific focus of his research is the problem of identifying sample-variable associations in large data sets that may arise from multiple measurement technologies. In the typical case where the data from an experiment are represented in the form of a rectangular matrix, sample-variable associations correspond to distinguished submatrices of the data matrix. The investigator is developing a statistically principled, significance-based approach to the problem of finding large average submatrices of a data matrix, using a simple iterative algorithm. The algorithm is applicable to real-valued and categorical data matrices. In addition to the basic method, the investigator is developing several extensions, including data-driven null models that incorporate dependence between variables, data arising from the simultaneous application of multiple measurement technologies, and application of the basic method to prediction problems such as classification, regression and survival analysis. In addition, the investigator is developing basic theory to support the use of the algorithm, and to assess the structure of data matrices under the different null models. The development and application of the methods is being carried out in close collaboration with several groups of biomedical researchers. In particular, the new data mining methodology is being incorporated into software that is used by collaborating scientists to identify and assess significant sample-variable associations in ongoing experiments involving breast, brain and lung cancer.Large data sets are now common in many experimental areas of science, and in particular gene-level studies of human diseases such as cancer. In such studies it is not unusual to encounter experiments containing from hundreds to thousands of samples, and tens of thousands to millions of measurements on each sample. Large data sets are part of a trend away from traditional hypothesis-driven scientific research towards data-driven research, in which researchers explore large data sets for patterns or regularities that, in conjunction with subject matter expertise, yield hypotheses that can be tested by more traditional means. The investigator is studying an exploratory method that identifies statistically significant associations between samples and variables in large data sets, associations that can yield testable scientific hypotheses. The methods being developed by the investigator are computationally efficient, and are based on established statistical principles, in particular the notion of statistical significance. The investigator is also studying ways in which the basic exploratory method can be applied to data arising from multiple measurement technologies, and application of the basic method to statistical problems such as classification and survival analysis. These activities are being carried out as part of a collaborative research program involving the sustained interactions of faculty and students from the statistical, biological, and medical sciences. The exploratory method developed by the investigator is being integrated into the basic exploratory tools of the collaborating scientists, and is a component in the analysis of several new, previously unanalyzed, data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Inference for Stationary Processes: Optimal Transport and Generalized Bayesian Approaches
Iterative testing procedures and high-dimensional scaling limits of extremal random structures
Optimality Landscapes and Exploratory Data Analysis
Analysis of High Dimensional Data Using Subspace Clustering
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    YU BYUNGJUN
  • 依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
  • 批准号:
    W2433169
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    HAOFEI ZHANG
  • 依托单位:
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    20万元
  • 批准年份:
    2020
  • 负责人:
    SAGAR RIZWAN UR REHMAN
  • 依托单位: