Adaptive Reproducible High-Dimensional Nonlinear Inference for Big Biological Data
Adaptive Reproducible High-Dimensional Nonlinear Inference for Big Biological Data
批准号:
9753295
负责人:
Yingying Fan
金额:
$27.99万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2022-04-30
关键词:
AddressAlgorithmsArchaeaAttentionBacteriaBig DataBiologicalBypassCellsColorectal CancerComplexComputer softwareConsultCoupledDataData SetDevelopmentDimensionsDiseaseEcosystemEffectivenessEnvironmentFoundationsFrequenciesGaussian modelGenesGenetic MaterialsGenomicsHealthcareHumanInternetInvestigationJointsLengthLinear RegressionsLiteratureLiver CirrhosisMachine LearningMarinesMathematicsMetagenomicsMethodsModelingModernizationMolecularMolecular Sequence DataMutationNeurosciencesNon-Insulin-Dependent Diabetes MellitusNon-linear ModelsObesityOrganismPerformancePlanet EarthPlayProceduresReproducibilityReproducibility of ResultsResearchResearch PersonnelRoleSamplingSampling StudiesShotgunsSocial SciencesTestingTheoretical StudiesTissuesTrainingViralVirusVisualization softwareWorkbasebiological researchcomputerized toolscontigdark matterdeep learningdeep learning algorithmdesignflexibilityhigh dimensionalityhuman diseasehuman tissueimprovedinterestlearning strategymetagenomic sequencingmicrobial communitymicrobiomemicrobiome researchmodel designmodel developmentnew technologynovelpower analysisresponsesimulationtheoriestraituser-friendlyvirus host interactionvirus identification
中文摘要
大数据现在无处不在,存在于现代科学研究的各个领域。许多当代应用,
例如最近的国家微生物组倡议(NMI),极大地需要高度灵活的统计机器
既能产生可解释的结果,又能产生可重复的结果的学习方法。因此,这是至关重要的。
重要的是要找出导致大量患者做出反应的关键原因。
可用协变量,可以统计地表示为中的错误发现率(FDR)控制
一般高维非线性模型。尽管鸟枪式元基因组学有着巨大的应用
研究方面,现有的大多数调查都集中在细菌生物体的研究上。然而,病毒
而病毒与宿主的相互作用在控制微生物群落的功能方面起着重要作用。在……里面
此外,病毒已被证明与复杂的疾病有关。然而,对这一事件的调查
病毒在人类疾病中的作用严重不足。这项建议的目的是
开发数学严谨和计算高效的方法来处理高度复杂的大
数据和这些方法的应用,以解决基本和重要的生物学和
生物医学问题。有四个相互关联的目标。在目标1中,我们将从理论上研究
最近提出的无模型仿冒(MFK)程序,从理论上证明了
控制任意型号和任意尺寸的FDR。我们还将从理论上证明健壮性
关于协变量分布的错误指定的MFK。这些研究将奠定基础
为了我们在其他目标上的发展。在目标2中,我们将开发深度学习方法来预测病毒
以更高的精度,将新算法与MFK相结合,实现了对病毒基元的FDR控制
发现,并调查我们的新程序的威力和健壮性。在目标3中,我们将考虑
考虑病毒-宿主基序的相互作用,并采用我们在目标2中的算法和理论来预测
病毒与宿主的感染相互作用状态。在目标4中,我们将应用前三种开发的方法
旨在分析实验中心的鸟枪式元基因组数据集以识别病毒和病毒宿主
在某些目标FDR水平上与几种疾病相关的相互作用。无论是算法还是结果
将通过网络传播。这项研究的结果将对元基因组学具有重要意义
在各种环境下学习。
英文摘要
Big data is now ubiquitous in every field of modern scientific research. Many contemporary applications,
such as the recent national microbiome initiative (NMI), greatly demand highly flexible statistical machine
learning methods that can produce both interpretable and reproducible results. Thus, it is of paramount
importance to identify crucial causal factors that are responsible for the response from a large number of
available covariates, which can be statistically formulated as the false discovery rate (FDR) control in
general high-dimensional nonlinear models. Despite the enormous applications of shotgun metagenomic
studies, most existing investigations concentrate on the study of bacterial organisms. However, viruses
and virus-host interactions play important roles in controlling the functions of the microbial communities. In
addition, viruses have been shown to be associated with complex diseases. Yet, investigations into the
roles of viruses in human diseases are significantly underdeveloped. The objective of this proposal is to
develop mathematically rigorous and computationally efficient approaches to deal with highly complex big
data and the applications of these approaches to solve fundamental and important biological and
biomedical problems. There are four interrelated aims. In Aim 1, we will theoretically investigate the power
of the recently proposed model-free knockoffs (MFK) procedure, which has been theoretically justified to
control FDR in arbitrary models and arbitrary dimensions. We will also theoretically justify the robustness
of MFK with respect to the misspecification of covariate distribution. These studies will lay the foundations
for our developments in other aims. In Aim 2, we will develop deep learning approaches to predict viral
contigs with higher accuracy, integrate our new algorithm with MFK to achieve FDR control for virus motif
discovery, and investigate the power and robustness of our new procedure. In Aim 3, we will take into
account the virus-host motif interactions and adapt our algorithms and theories in Aim 2 for predicting
virus-host infectious interaction status. In Aim 4, we will apply the developed methods from the first three
aims to analyze the shotgun metagenomics data sets in ExperimentHub to identify viruses and virus-host
interactions associated with several diseases at some target FDR level. Both the algorithms and results
will be disseminated through the web. The results from this study will be important for metagenomics
studies under a variety of environments.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adaptive Reproducible High-Dimensional Nonlinear Inference for Big Biological Data
-
批准号:9674585
-
项目类别:
-
资助金额:$28.97万
-
财政年份:2018
-
负责人:Yingying Fan
-
依托单位:
Adaptive Reproducible High-Dimensional Nonlinear Inference for Big Biological Data
-
批准号:10159277
-
项目类别:
-
资助金额:$27.67万
-
财政年份:2018
-
负责人:Yingying Fan
-
依托单位:
Adaptive Reproducible High-Dimensional Nonlinear Inference for Big Biological Data
-
批准号:9923688
-
项目类别:
-
资助金额:$27.67万
-
财政年份:2018
-
负责人:Yingying Fan
-
依托单位:
海外基金