课题基金 / 基金详情

Constrained Group Selection and Structure Estimation in Semiparametric Models

Constrained Group Selection and Structure Estimation in Semiparametric Models
半参数模型中的约束组选择和结构估计
批准号:
1208225
负责人:
Jian Huang
金额:
$15.97万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-07-01 至 2015-06-30

项目摘要

项目成果

Jian Huang的其他基金

相似基金

相关文献

中文摘要
翻译
在高维模型中,当参数存在自然约束时,提出了一类新的约束群选择方法。该项目有望为研究几个重要的统计建模和分析问题提供新的研究方向,这些问题包括半参数加性模型、变系数模型中的结构估计和变量选择以及变量数量大于样本量的高维环境下的生存分析模型。拟议的项目还将产生多个基因组数据集的综合分析和全基因组关联研究的新方法。所提出的方法在高维环境下的理论性质和计算算法将被开发。高维数据的分析在统计学中提出了新的和具有挑战性的理论和计算问题。假设变量数目是固定的且比样本量小得多的标准方法不适用于高维模型。所提出的方法有望在稀疏、高维的环境中正确地选择重要的组并正确地估计高概率的模型结构。高维数据出现在许多不同的科学和人文领域,包括生物、经济、金融、信息技术和健康科学。在所有这些领域中,特征选择是从数据中发现知识的关键步骤。在遗传和基因组研究方面,随着生物技术的快速发展,产生了越来越多的大数据集。从高维和噪声数据集中识别具有统计学和生物学意义的模式正成为一项重大挑战。在估计临床结果和遗传数据之间的关系方面,能够处理高维问题的统计方法的发展将有助于更好地理解疾病的遗传基础,更好地诊断和更好地预测生存。该方法将应用于高维删失生存数据、纵向数据、全基因组关联研究和多基因组数据集的综合分析。在许多临床和生物医学研究中出现了删失和纵向数据。广义遗传多样性分析和综合分析是寻找常见和复杂疾病易感基因的重要方法。临床和遗传学研究的最终目标是了解危险因素和表型之间的关系,以开发新的疾病预防、诊断和治疗方法。该项目旨在将新的统计方法转化为分析高维临床和基因组数据的新方法,这些数据对实现这一目标非常重要。拟议项目的方法和结果将纳入关于高维数据分析的研究生课程。研究人员将通过向科学期刊提交论文并在互联网上公开提供这些论文和计算机程序,向科学界广泛传播研究结果。调查员还将在科学会议和研讨会上介绍研究结果。
英文摘要
This application proposes a class of novel constrained group selection methods in high-dimensional models when there are natural constraints on the parameters. The proposed project is expected to stimulate new research directions for studying several important statistical modeling and analysis problems, which include structure estimation and variable selection in semiparametric additive models, varying coefficient models, and survival analysis models in high-dimensional settings where the number of variables is larger than the sample size. The proposed project will also yield new methods for integrative analysis of multiple genomic datasets and genome wide association studies. Theoretical properties of the proposed methods in high-dimensional settings and computational algorithms will be developed. Analysis of high-dimensional data presents new and challenging theoretical and computational questions in statistics. Standard methods assuming the number of variables is fixed and much smaller than the sample size are not applicable to high-dimensional models. The proposed methods are expected to be able to correctly select the important groups and correctly estimate model structures with high probability in sparse, high-dimensional settings. High-dimensional data arise in many diverse fields of sciences and humanities, including biology, economics, finance, information technology, and health sciences. In all these fields, feature selection is a crucial step in the process of knowledge discovery from data. In genetic and genomic research, with rapid advances in biotechnology, more and more big data sets are being generated. The identification of statistically and biologically significant patterns from high-dimensional and noisy data sets is becoming a major challenge. The development of statistical methods that can deal with high-dimensional problems in estimating the relationship between clinical outcomes and genetic data will contribute to better understanding of the genetic basis of diseases, better diagnoses, and better survival prediction. The proposed methods will be applied to the analysis of high-dimensional censored survival data, longitudinal data, genome wide association studies (GWAS) and integrative analysis of multiple genomic datasets. Censored and longitudinal data arise in many clinical and biomedical studies. GWAS and integrative analysis are important methods for identifying disease susceptibility genes for common and complex diseases. The ultimate goal of clinical and genetic research is to understand the relationships between risk factors and phenotypes for developing new approaches to prevention, diagnosis and treatment of disease. This project aims to translate novel statistical approaches into new methodologies for analyzing high-dimensional clinical and genomic data that are important in achieving this goal. The methods and results from the proposed project will be incorporated into a graduate course on high-dimensional data analysis. The investigator will broadly disseminate the results to the scientific community by submitting papers to scientific journals and making them and the computer programs publicly available on the internet. The investigator will also present the results in scientific conferences and workshops.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Towards Learning-Based Storage Systems with Hardware-Software Co-Design
Collaborative Research: Elements: Towards A Scalable Infrastructure for Archival and Reproducible Scientific Visualizations
  • 批准号:
    2209767
  • 项目类别:
    Standard Grant
  • 资助金额:
    $31.62万
  • 财政年份:
    2022
  • 负责人:
    Jian Huang
  • 依托单位:
EAGER: CRYO: Continuous Adiabatic Demagnetization Refrigeration Below 1K without Helium-3
  • 批准号:
    2232489
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.97万
  • 财政年份:
    2022
  • 负责人:
    Jian Huang
  • 依托单位:
Collaborative Research: Integrating multi-dimensional omics data for quantifying disease heterogeneity
  • 批准号:
    1916199
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.03万
  • 财政年份:
    2019
  • 负责人:
    Jian Huang
  • 依托单位:
国内基金
海外基金
分泌蛋白IGFBP2在儿童Group3/Group4型髓母细胞瘤恶性进展中的作用与机制研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    夏明杨
  • 依托单位:
大兴安岭火山湖Group I长链烯酮冷季节温标研究与过去2000年温度定量重建
  • 批准号:
    42073070
  • 项目类别:
    面上项目
  • 资助金额:
    61.0万元
  • 批准年份:
    2020
  • 负责人:
    姚远
  • 依托单位:
近海沉积物中Marine Group I古菌新类群的发现、培养及其驱动碳氮循环的机制
  • 批准号:
    92051115
  • 项目类别:
    重大研究计划
  • 资助金额:
    81.0万元
  • 批准年份:
    2020
  • 负责人:
    刘吉文
  • 依托单位:
MicroRNA靶向的漆酶基因及其所在Group 1 亚家族成员 调控水稻产量性状的功能机制
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    257万元
  • 批准年份:
    2019
  • 负责人:
    陈月琴
  • 依托单位: