基于遗传变异的跨组学整合统计方法及其在肺癌风险中的研究
批准号:
82103946
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
沈思鹏
依托单位:
学科分类:
流行病学方法与卫生统计
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
沈思鹏
中文摘要
识别人群的肿瘤易感风险对于疾病的一级预防具有十分重要的意义。整合多个组学的生物标志物可以加深对肿瘤的认知,提高预测精度。本课题基于肺癌大规模的遗传变异数据,进行跨组学整合的统计方法学研究。通过整合多种数量性状基因座(QTL)数据库效应,使用基于贝叶斯的统计学方法估计基因表达、代谢物和DNA甲基化水平;通过多组学肿瘤风险关联研究(xWAS),筛选出于肺癌发病风险关联的生物标志物,进而从多个层面解释预测疾病;最后基于多组学数据使用机器学习与惩罚回归相结合的多阶段统计策略构建肺癌风险预测模型,提高模型稳健性与预测精度。本课题的顺利实施,将为从血液遗传变异预测人群的肿瘤发病风险提供全新方法学和应用成果,将从多组学角度结合各组学的优势,兼顾样本量和结果可解释性,也将为复杂疾病的精准风险预测提供强有力的手段,具有重要的科学意义和实用价值。
英文摘要
It is very important to identify the risk of cancer susceptibility for primary prevention of diseases. Integrating multi-omics biomarkers can deepen the understanding of tumor and improve the prediction accuracy. As the sample size of genetic variation data is more than other omics, it is worth developing statistical methods and make full use of genetic variation data to deeply mine the information of multi-omics data and explore their association with diseases. Based on the large-scale genetic variation data of lung cancer, this project carries out the statistical methodology research of cross-omics integration. By integrating the effects of multiple quantitative trait loci (QTL) database, we develop the statistical methods based on Bayesian to estimate gene expression, metabolites and DNA methylation levels. By using the multi-omics association study (xWAS), we screen the biomarkers related to the risk of lung cancer, and then explained and predicted the disease from multiple levels. Finally, we predict the risk of lung cancer based on the multi-omics data. A multi-stage statistical strategy combining with machine learning and penalty regression is used to develop a lung cancer risk prediction model to improve the robustness and prediction accuracy. The successful implementation of this project will provide a new methodology and application example for predicting the risk of lung cancer from blood genetic variation, which has important scientific significance.
课题组聚焦重大复杂肿瘤的精准防控,采用统计遗传学、分子流行病学、生物信息学等多学科交叉手段,围绕大型人群队列的多种组学数据开展研究,系统研究了肿瘤病因预防、肿瘤筛查预防、肿瘤预后改善潜在生物标志物和预防措施。具体来说,研究主要集中在以下三个方面:(1)基于大型人群基因组学数据,进行风险因素识别和跨组学推断;(2)大型肿瘤人群队列多组学整合研究;(3)基于大型人群宏观因素和组学数据构建肿瘤预测模型。在进行人群队列多组学研究等关键科学技术难点上取得了一系列原创成果,发表Cell Reports, Am J Respir Crit Care Med, J Thorac Oncol等多篇中科院一区TOP权威期刊,并被多次引用和专题报告,为肿瘤的精准防控提供了重要科学依据。已有的研究工作为进一步开展中国超大人群队列测序等多组学研究,发现肿瘤预防干预靶点和优化风险预测模型奠定了扎实的理论基础。
基于超大人群全基因组测序的肺癌胚系突变关联整合方法研究
-
批准号:82373685
-
项目类别:面上项目
-
资助金额:49万元
-
批准年份:2023
-
负责人:沈思鹏
-
依托单位:
国内基金
海外基金