Rule-based machine learning to address heterogeneity in high-dimensional survival data
Rule-based machine learning to address heterogeneity in high-dimensional survival data
批准号:
10478828
负责人:
Alexa Abigail Woodward
金额:
$2.51万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2022-12-05
关键词:
AddressAdultAgeAlgorithmsAlkylating AgentsArchitectureAutomobile DrivingBiologicalBrain NeoplasmsCancer PrognosisCellsClinicalComplexCox ModelsCpG Island Methylator PhenotypeDNADNA IntegrationDNA MethylationDNA SequenceDNA Sequence AlterationDataData SetDevelopmentDimensionsDiseaseDisease OutcomeDrug TargetingElementsEpigenetic ProcessFoundationsGenesGeneticGenetic EpistasisGenetic HeterogeneityGenetic ModelsGenetic TranscriptionGenomicsGliomaHeadHeritabilityHeterogeneityHistonesHypermethylationIncidenceInformaticsInterventionMGMT geneMachine LearningMalignant - descriptorMalignant NeoplasmsMethodologyMethodsMethylationMicroRNAsModelingModificationMolecularMultiomic DataOutcomeOutputPathway AnalysisPathway interactionsPatternPerformancePersonsPharmacologyPhenotypePlayPrecision therapeuticsPrediction of Response to TherapyPrimary Brain NeoplasmsPrognosisPromoter RegionsProteomicsRNAResearchResearch PersonnelRiskSample SizeSingle Nucleotide PolymorphismSomatic MutationSystemTestingThe Cancer Genome AtlasTrainingTreatment EfficacyUpdateValidationVisualVisualizationVisualization softwarebasecancer heterogeneitycancer riskcancer typedata modelingdeep learningdisorder riskepigenomicsexperiencefeature selectionforestgenetic analysisgenetic architecturegenetic epidemiologygenome wide association studygenome wide methylationgenomic datahigh dimensionalityimprovedinnovationinsightlarge datasetslearning classifiermachine learning methodmultidimensional datamultiple omicsnovelpersonalized approachpersonalized carepotential biomarkerpreservationpromoterskillssuccesssurvival predictiontemozolomidetherapeutic targettherapy resistanttooltranscriptomicstreatment responsetreatment strategytumortumor progression
中文摘要
项目摘要
在后基因组时代,研究人员面临着大量的数据需要分析和解释。基因组-
广谱关联分析通常拥有数百万个单核苷酸多态(SNPs)以及
越来越大的表观基因组、转录组、蛋白质组(多组体)和其他数据集。而当前的
遗传流行病学的标准强调增加样本量,我们建议取得实质性进展
可以通过开发改进的方法来分析海量的多组数据,目前
是存在的。包括维度和多重测试负担在内的一些方法学挑战
到目前为止,许多方法的成功受到了限制。此外,只考虑简单的线性关联
忽略了更有可能的情况,即复杂的遗传和多经济关系推动风险和结果
在常见疾病中。异质性只是导致疾病风险和
结果,但可以说是最难建模和检测的。这个项目解决了这个问题和其他问题
挑战胶质瘤,这是一种高度异质性的癌症类型。改进现有的治疗策略
癌症和胶质瘤无疑需要对遗传异质性和
表观遗传机制。除了使用一种
特征选择策略,可以同时检测主效应和交互作用,并保持异质性,我们将
修改现有的异质性检测方法,以适应经审查的生存数据。首先,在目标1中,
我们将使用模拟的遗传生存数据来建立基于RELAGE的特征选择算法的实用性
在捕捉复杂的遗传结构(即主效应、异质性和上位性)方面。我们会比较一下
针对生存数据的高维特征选择的标准方法。AIM 2更新学习
分类器系统(LCS),一种基于规则的机器学习,它使用IF/THEN规则来建模复杂的
异类问题空间。据我们所知,还没有处理经过审查的生存数据的LC
发展到目前为止。在模拟数据上测试我们的生存LCS并将其与标准生存进行比较后
方法,在目标3中,我们将使用TCGA胶质瘤的体细胞突变和甲基化数据来实现它
数据集。最后,作为目标3的一部分,我们将使用LCS输出执行路径分析,以努力
确定异质关联背后的共同生物途径。我们还将利用一个网络
可视化工具,可更好地了解要素之间的交互,并提供对
结果。该项目的发现将为胶质瘤的精准护理和治疗奠定基础。我们的
高维、异质生存数据的创新方法将既可推广又可
可解释的,当前机器学习方法所缺少的品质。这个项目和
不可否认,随附的培训计划为培养必要的技能和经验提供了一个理想的环境
成为遗传流行病学和信息学前沿的独立调查者。
英文摘要
Project Summary
In the post-genomic era, researchers are met with an abundance of data to analyze and interpret. Genome-
wide association analyses (GWAS) often boast millions of single-nucleotide polymorphisms (SNPs), alongside
increasingly large epigenomic, transcriptomic, proteomic (multi-omic) and other data sets. While the current
standard in genetic epidemiology emphasizes increased sample sizes, we propose that substantial progress
can be made by developing improved methods to analyze the vast amount of multi-omic data that currently
exists. A number of methodological challenges including dimensionality and the multiple testing burden have
limited the success of many approaches thus far. Furthermore, only considering simple, linear associations
leaves out the more likely scenario of complex genetic and multi-omic relationships driving risk and outcomes
in common diseases. Heterogeneity is just one of the complex mechanisms that underlies disease risk and
outcomes, but is arguably among the most difficult to model and detect. This project tackles this and other
challenges in glioma, a highly heterogeneous cancer type. Improving upon available treatment strategies in
cancer and glioma specifically will undoubtedly require a full characterization of genetic heterogeneity and
epigenetic mechanisms. In addition to confronting the dimensionality of genetic and epigenetic data using a
feature selection strategy that can detect both main effects and interaction and preserve heterogeneity, we will
modify an existing method for detecting heterogeneity to accommodate censored survival data. First, in Aim 1,
we will use simulated genetic survival data to establish the utility of a Relief-based feature selection algorithm
in capturing complex genetic architectures (i.e., main effects, heterogeneity, and epistasis). We will compare it
against standard approaches for high-dimensional feature selection of survival data. Aim 2 updates a learning
classifier system (LCS), a type of rule-based machine learning that uses IF/THEN rules to model complex and
heterogeneous problem spaces. To our knowledge, no LCS that handles censored survival data has been
developed to date. After testing our survival LCS on simulated data and comparing it to standard survival
methods, in Aim 3 we will implement it using somatic mutation and methylation data from the TCGA glioma
dataset. Finally, as part of Aim 3, we will perform a pathway analysis using the LCS output in an effort to
identify common biological pathways underlying heterogeneous associations. We will also utilize a network
visualization tool to better understand interactions between features and provide a visual interpretation of the
results. Findings from this project will lay the foundation for precision care and treatment of glioma. Our
innovative approach to high-dimensional, heterogeneous survival data will be both generalizable and
interpretable, qualities that are missing from current machine learning approaches. This project and the
accompanying training plan undeniably provide an ideal setting to develop the skills and experience necessary
to become and independent investigator at the forefront of genetic epidemiology and informatics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金