Statistical methods for ovarian cancer diagnosis and prognosis
Statistical methods for ovarian cancer diagnosis and prognosis
批准号:
2632835
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
高通量测序技术的发展导致了大规模分析数据的产生;使我们能够深入了解潜在的生物过程。在不同的层次上,测序使我们能够收集有关DNA、RNA、蛋白质、代谢物等的数据,在描述生物对象时提供补充信息。单独的每个数据源,被称为组学数据,表征了生物体的特定部分。例如,基于DNA水平的基因组学表征了基因组,而基于RNA水平的转录组学表征了转录组。值得注意的是,每个水平都是相互关联的,例如,mRNA被翻译成蛋白质,驱动细胞的行为,从而导致表型的表达。由于组学数据集的高维性和异质性,该分析面临着统计方面的挑战。在我的博士学位期间,我将致力于研究新的统计方法,解决组学数据集的降维和变量选择问题,包括监督和无监督设置。目前正在开发的一种方法是高维贝叶斯生存分析模型,该模型使用尖钉-板先验。我们的方法使我们能够在高维环境中进行变量选择,同时也提供了不确定性量化和效果估计的机制。在生物医学科学中,生存分析是一项至关重要的任务,当使用转录组学数据进行分析时,可以创建预后模型和发现生物标志物。我博士学位的第二个方面将集中在数据集成方法的发展上。其中数据集成涉及多个数据集的联合分析,目的是理解它们之间的关系。在帝国理工大学CRUK中心的合作者的推动下,我们将把这些方法应用于放射组学数据(从医学图像构建的图像特征),以及从卵巢癌患者收集的其他组学数据集。因此,为经济实惠且易于收集的(CT/MRI)扫描提供生物学解释。目前,我们正在考虑扩展典型相关分析的概率框架。这种扩展将使这些方法能够在高维环境中工作,同时提供不确定性量化。与EPSRC在人工智能和医疗保健方面的战略保持一致,拟议的方法发展旨在通过优化患者治疗来改善医疗服务。最终,我博士的重点是基于卵巢癌患者的数据,然而生物学相关方法的一般适用性超出了单一疾病。
英文摘要
The development of high-throughput sequencing technologies has led to the production of large-scale profiling data; allowing us to gain insight into underlying biological processes. Available at different levels, sequencing allows us to collect data about DNA, RNA, proteins, metabolites and so forth, providing complementary information when characterising a biological object. Individually each source of data, referred to as omics data, characterises a specific part of an organism. For instance, genomics based at the DNA level characterises the genome, whilst transcriptomics, based at the RNA level characterises the transcriptome. Notably, each level is related to one another, for instance, mRNA is translated to proteins, driving the behaviour of cells thereby leading to the expression of phenotypes. Due to the high-dimensionality and heterogeneity within omics datasets, the analysis is ripe with statistical challenges. Throughout my PhD I will be working on novel statistical methods tackling the issues of dimensionality reduction and variable selection for omics datasets, both in supervised and unsupervised settings. One such method, currently under development, is a high-dimensional Bayesian survival analysis model that uses a spike-and-slab prior. Our method enables us to perform variable selection in a high-dimensional setting, whilst also offering mechanisms for uncertainty quantification and effect estimation. Within the biomedical sciences, survival analysis is a task of key importance, and when performed with transcriptomics data enables the creation of prognostic model and the discovery of biomarkers. A second aspect of my PhD will focus on the development of methodology for data-integration. Where data-integration involves the joint analysis of multiple datasets with the goal of understanding the relationships between them. Motivated by our collaborators at Imperial's CRUK centre, we will be applying these methods to radiomics data (image features constructed from medical images), and other omics datasets collected from patients with ovarian cancer. Thereby, providing biological interpretations to affordable and easy to collect (CT/MRI) scans. Currently, we are considering extending the probabilistic framing of canonical correlation analysis. Such extensions will enable these methods to work in a high-dimensional setting and simultaneously provide uncertainty quantification. Aligning with EPSRC strategies in artificial intelligence and healthcare, the proposed methodological developments seek to improve health services by optimising patient treatments. Ultimately, the focus of my PhD is based on data from patients with ovarian cancer, however the general applicability of biologically relevant methods extends beyond single disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
复杂图像处理中的自由非连续问题及其水平集方法研究
-
批准号:60872130
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2008
-
负责人:刘国才
-
依托单位:
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: