Rule-based machine learning to address heterogeneity in high-dimensional survival data
Rule-based machine learning to address heterogeneity in high-dimensional survival data
批准号:
10478828
负责人:
Alexa Abigail Woodward
金额:
$2.51万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2022-12-05
关键词:
AddressAdultAgeAlgorithmsAlkylating AgentsArchitectureAutomobile DrivingBiologicalBrain NeoplasmsCancer PrognosisCellsClinicalComplexCox ModelsCpG Island Methylator PhenotypeDNADNA IntegrationDNA MethylationDNA SequenceDNA Sequence AlterationDataData SetDevelopmentDimensionsDiseaseDisease OutcomeDrug TargetingElementsEpigenetic ProcessFoundationsGenesGeneticGenetic EpistasisGenetic HeterogeneityGenetic ModelsGenetic TranscriptionGenomicsGliomaHeadHeritabilityHeterogeneityHistonesHypermethylationIncidenceInformaticsInterventionMGMT geneMachine LearningMalignant - descriptorMalignant NeoplasmsMethodologyMethodsMethylationMicroRNAsModelingModificationMolecularMultiomic DataOutcomeOutputPathway AnalysisPathway interactionsPatternPerformancePersonsPharmacologyPhenotypePlayPrecision therapeuticsPrediction of Response to TherapyPrimary Brain NeoplasmsPrognosisPromoter RegionsProteomicsRNAResearchResearch PersonnelRiskSample SizeSingle Nucleotide PolymorphismSomatic MutationSystemTestingThe Cancer Genome AtlasTrainingTreatment EfficacyUpdateValidationVisualVisualizationVisualization softwarebasecancer heterogeneitycancer riskcancer typedata modelingdeep learningdisorder riskepigenomicsexperiencefeature selectionforestgenetic analysisgenetic architecturegenetic epidemiologygenome wide association studygenome wide methylationgenomic datahigh dimensionalityimprovedinnovationinsightlarge datasetslearning classifiermachine learning methodmultidimensional datamultiple omicsnovelpersonalized approachpersonalized carepotential biomarkerpreservationpromoterskillssuccesssurvival predictiontemozolomidetherapeutic targettherapy resistanttooltranscriptomicstreatment responsetreatment strategytumortumor progression
中文摘要
项目摘要
在后基因组时代,研究人员遇到了大量的数据进行分析和解释。基因组-
广泛关联分析(GWAS)通常拥有数百万个单核苷酸多态性(SNP),
越来越大的表观基因组、转录组、蛋白质组(多组)和其他数据集。虽然目前的
遗传流行病学标准强调增加样本量,我们建议,
可以通过开发改进的方法来分析大量的多组学数据,
存在.包括维度和多重测试负担在内的一些方法学挑战,
限制了迄今为止许多方法的成功。此外,仅考虑简单的线性关联,
忽略了更可能的情况,即复杂的遗传和多组学关系驱动风险和结果
在常见疾病中。异质结只是疾病风险的复杂机制之一,
结果,但可以说是最难建模和检测的。这个项目解决了这个和其他
神经胶质瘤是一种高度异质性的癌症类型。改善现有的治疗策略,
癌症和神经胶质瘤无疑需要对遗传异质性进行全面表征,
表观遗传机制除了使用一种新的方法来面对遗传和表观遗传数据的维度之外,
特征选择策略,可以检测主效应和交互作用,并保持异质性,我们将
修改现有的异质性检测方法,以适应删失生存数据。首先,在目标1中,
我们将使用模拟的遗传生存数据来建立一种基于救济的特征选择算法的效用
在捕获复杂的遗传结构(即,主效应、异质性和上位性)。我们会比较一下
针对生存数据的高维特征选择的标准方法。目标2更新学习
分类器系统(LCS),一种基于规则的机器学习,使用IF/THEN规则对复杂的
异构问题空间据我们所知,没有处理删失生存数据的LCS,
发展至今。在模拟数据上测试我们的生存LCS并将其与标准生存进行比较后,
方法,在目标3中,我们将使用来自TCGA胶质瘤的体细胞突变和甲基化数据来实现它。
数据集。最后,作为目标3的一部分,我们将使用LCS输出进行路径分析,
确定异质关联背后的共同生物学途径。我们还将利用网络
可视化工具,以更好地了解功能之间的相互作用,并提供可视化的解释
结果该项目的发现将为胶质瘤的精确护理和治疗奠定基础。我们
高维、异质生存数据创新方法将是可推广的,
可解释的,当前机器学习方法所缺少的品质。本工程与
随附的培训计划为培养必要的技能和经验提供了理想的环境
成为遗传流行病学和信息学前沿的独立研究者。
英文摘要
Project Summary
In the post-genomic era, researchers are met with an abundance of data to analyze and interpret. Genome-
wide association analyses (GWAS) often boast millions of single-nucleotide polymorphisms (SNPs), alongside
increasingly large epigenomic, transcriptomic, proteomic (multi-omic) and other data sets. While the current
standard in genetic epidemiology emphasizes increased sample sizes, we propose that substantial progress
can be made by developing improved methods to analyze the vast amount of multi-omic data that currently
exists. A number of methodological challenges including dimensionality and the multiple testing burden have
limited the success of many approaches thus far. Furthermore, only considering simple, linear associations
leaves out the more likely scenario of complex genetic and multi-omic relationships driving risk and outcomes
in common diseases. Heterogeneity is just one of the complex mechanisms that underlies disease risk and
outcomes, but is arguably among the most difficult to model and detect. This project tackles this and other
challenges in glioma, a highly heterogeneous cancer type. Improving upon available treatment strategies in
cancer and glioma specifically will undoubtedly require a full characterization of genetic heterogeneity and
epigenetic mechanisms. In addition to confronting the dimensionality of genetic and epigenetic data using a
feature selection strategy that can detect both main effects and interaction and preserve heterogeneity, we will
modify an existing method for detecting heterogeneity to accommodate censored survival data. First, in Aim 1,
we will use simulated genetic survival data to establish the utility of a Relief-based feature selection algorithm
in capturing complex genetic architectures (i.e., main effects, heterogeneity, and epistasis). We will compare it
against standard approaches for high-dimensional feature selection of survival data. Aim 2 updates a learning
classifier system (LCS), a type of rule-based machine learning that uses IF/THEN rules to model complex and
heterogeneous problem spaces. To our knowledge, no LCS that handles censored survival data has been
developed to date. After testing our survival LCS on simulated data and comparing it to standard survival
methods, in Aim 3 we will implement it using somatic mutation and methylation data from the TCGA glioma
dataset. Finally, as part of Aim 3, we will perform a pathway analysis using the LCS output in an effort to
identify common biological pathways underlying heterogeneous associations. We will also utilize a network
visualization tool to better understand interactions between features and provide a visual interpretation of the
results. Findings from this project will lay the foundation for precision care and treatment of glioma. Our
innovative approach to high-dimensional, heterogeneous survival data will be both generalizable and
interpretable, qualities that are missing from current machine learning approaches. This project and the
accompanying training plan undeniably provide an ideal setting to develop the skills and experience necessary
to become and independent investigator at the forefront of genetic epidemiology and informatics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金