Uncovering therapeutic-associated biomarkers via machine learning and feature engineering approaches
Uncovering therapeutic-associated biomarkers via machine learning and feature engineering approaches
批准号:
10564098
负责人:
Hu Li
金额:
$31.8万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-09-20 至 2024-09-19
关键词:
BiologicalBiological MarkersClassificationClinicalClinical TrialsCommunicationComputer AnalysisDataData SetDevelopmentDiagnosisDiagnosticDiagnostic Reagent KitsDiseaseDisease OutcomeDrug TargetingEncapsulatedEngineeringEtiologyExhibitsExpression LibraryFingerprintFundingGenesGeneticGenetic HeterogeneityGenetic VariationGenomeGenotypeGoalsHeterogeneityHomology ModelingIndividualLibrariesLightMachine LearningMalignant NeoplasmsMapsMedicineMethodsMolecularMolecular ProfilingNamesNetwork-basedPathologicPatientsPerformancePharmaceutical PreparationsPharmacologyPhenotypePlayProcessPropertyRNAResearchRoleSignal TransductionSpecificityTestingTherapeuticTimeTissue-Specific Gene ExpressionTissuesTranscriptTranslational ResearchUncertaintyUnited States National Institutes of HealthValidationWorkbasebiological systemsbiomarker discoverycancer typecandidate markercohortcomparativediagnostic valuedrug repurposingextracellularindividual patientindividual variationinnovationmolecular phenotypenovelpressureresponsetherapeutic targettraittranscriptomics
中文摘要
项目摘要
识别具有诊断性、稳健性和可在个体间推广的生物标志物,
治疗价值是医学上最需要的奋进。然而,在这方面存在着许多挑战。
鉴定这种稳健的治疗相关生物标志物(TAB)。例如,目前的大多数方法
试图在一般患者组群中获得统计学上显著的差异生物信号,但未能
承认个体患者之间的异质性遗传背景和表型多样性。我们最近
研究使用新开发的基于机器学习的特征工程方法,并在一个
跨12种癌症类型的泛癌症研究显示,生物学上受约束的特征(在此命名
不变特征)在疾病中是普遍的,并且可以用于对个体癌症进行分类。重要的是,我们还
表明不变特征可用于构建从头生物网络并发现网络枢纽,
可以成功地用于推断相关基因的表达。因此,不变特征可以充当
信息编码器利用来自药物再利用中心的信息,我们表明这些中心基因也是
药物靶点总的来说,这些观察结果表明,不变特征中心可以是TAB候选者。我们
我建议在生物学约束的新视角下,我们可以使用生物标志物的动态方法
这一发现涵盖了个体患者之间的遗传异质性和分子波动。
我们的中心假设是,疾病状态在其分子活动中显示出限制,并且可识别
不变的特征具有诊断和治疗价值。这项提案的主要目的是揭示
使用选定的NIH共同基金数据集(即exRNA,GTEx,LINC和IDG)的TAB。在目标1中,我们
测试生物约束不变特征对于大多数生物状态(如果不是所有生物状态)是通用的这一假设。
我们将通过从选定的共同基金中找到每个生物状态的不变特征来证明这一点
数据集。我们将在疾病和正常状态下进行比较分析,以剖析疾病特异性
不变特征接下来,在目标2中,我们将测试不变特征中心是TAB的假设。我们将展示
这是通过确定不变特征中心的“可编码性”的诊断能力来重建
在不同的个体患者中,其相关的不变特征基因的表达值被诊断为
相同的疾病类型最后,我们将这些不变特征中心映射到IDG和DrugBank,以确定它们的
可用药性对于那些没有已知药物的未充分研究的枢纽,我们将进行计算分析,例如
同源建模和机器学习来表征它们的可药用性。我们期待及时完成
建议的目标和成功完成这一项目无疑将为选定的增值
共同基金数据集,同时提供了生物标志物和治疗靶点发现的新范式转变。
英文摘要
PROJECT SUMMARY
Identifying biomarkers that are diagnostic, robust and generalizable across individuals while possessing
therapeutic values is the most wanted endeavor in medicine. However, there are numerous challenges in the
identification of such robust therapy-associated biomarkers (TABs). For example, most of the current methods
seek to achieve statistically significant differential biological signals in general patient cohorts but fail to
acknowledge heterogenous genetic backgrounds and phenotypic diversity among individual patients. Our recent
studies using newly developed machine learning-based feature engineering approaches and conducted in a
pan-cancer study across 12 cancer types showed that biologically constrained features (named herein
invariant features) are universal in disease and can be used to classify individual cancers. Importantly, we also
show that invariant features can be used to build de novo biological networks and discover network hubs that
can be successfully utilized to infer the expression of associated genes. As such, invariant features can act as
information encoders. Using information from Drug Repurposing Hub we show that these hub genes are also
drug targets. Collectively, these observations suggest that invariant feature hubs can be TAB candidates. We
propose that under the new light of biological constraints, we can use a dynamic approach for biomarker
discovery that encapsulates both the genetic heterogeneity and molecular fluctuation across individual patients.
Our central hypothesis is that disease states show constrains in their molecular activities, and identifiable
invariable features possess diagnostic and therapeutic values. The main objective of this proposal is to uncover
TABs using selected NIH Common Fund datasets (namely, exRNA, GTEx, LINC, and IDG). In Aim 1, we will
test the hypothesis that biologically constrained invariant features are universal to most if not all biological states.
We will show this by finding invariant features with respect to each biological state from selected Common Fund
datasets. We will conduct comparative analyses in disease and normal states in order to dissect disease-specific
invariant features. Next, in Aim 2, we will test the hypothesis that invariant feature hubs are TABs. We will show
this by determining the diagnostic capability of invariant feature hubs for their “encodability” to reconstruct the
expression values of their associated invariant feature genes in different individual patients diagnosed under
same disease type. Finally, we will map these invariant feature hubs to IDG and DrugBank to determine their
druggability. For those understudied hubs with no known drugs, we will perform computational analyses such as
homology modeling and machine learning to characterize their druggability. We expect timely accomplishment
of proposed aims and successful completion of this project will no doubt provide added values for the selected
Common Fund datasets, while providing a new paradigm shift of biomarker and therapeutic target discovery.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Capturing the molecular complexity of Alzheimer's disease through the lens of RNA binding proteins
-
批准号:10249415
-
项目类别:
-
资助金额:$41.09万
-
财政年份:2018
-
负责人:Hu Li
-
依托单位:
INHA WITH INHIBITORS
-
批准号:8363387
-
项目类别:
-
资助金额:$0.19万
-
财政年份:2011
-
负责人:Hu Li
-
依托单位:
海外基金