课题基金 / 基金详情

Interpretable statistical machine learning approaches for the molecular investigation of cancer

Interpretable statistical machine learning approaches for the molecular investigation of cancer
用于癌症分子研究的可解释统计机器学习方法
批准号:
2728935
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
卵巢癌是英国女性第六常见的癌症。高级别浆液性卵巢癌(HGSOC)占大多数病例,5年生存率低至30%。导致这种不良预后的两个主要因素是:1)疾病的诊断较晚,以及2)尽管对治疗有初步反应,但复发的比例很高。后者表明,可能存在小群体的治疗抗性癌细胞,可以重新填充疾病。因此,识别这些癌细胞并了解它们与其他癌细胞类型的不同之处是很有意义的,这些癌细胞类型也可能影响患者对治疗的反应不同以及存活时间不同的原因。卵巢癌的分子基础可以使用大量的现代技术来解开,例如在大块组织和单细胞水平上的测序和成像。这创造了一个前所未有的机会,可以使用数据驱动的方法来精确描述卵巢癌的特征,并有可能开发有针对性的治疗方案。然而,为了有效地利用分子数据,需要稳健的分析方法来鉴定感兴趣的细胞群(特别是罕见的)和整合异构数据模式。这项研究的目的是:1)开发一种稳健且可解释的方法来从高维分子数据中识别稀有细胞群,(二)建立一个统计学框架,用于整合多模态数据进行生存预测。研究方法的新奇我们将开发专门用于从分子水平识别罕见细胞类型的统计技术。数据经典的统计发现方法往往偏向于最常见的细胞群,因为它们有更多的信息。通常存在与提出罕见细胞类型相关的惩罚,因为这些可能不是真实的,因此必须在提出新细胞类型和这些建议在进一步检查后可能被证明是错误的机会之间取得平衡。我们将开发技术,使我们能够控制这些相互竞争的需求之间的平衡,使实验科学家能够根据他们可接受的风险水平调整预期。我们还将开发技术,以结合联合收割机不同来源的数据,如临床记录,磁共振成像和全基因组测序。这些技术将检查现有方法的特定限制,这些方法通常不考虑不同类型数据之间的信息不平衡。例如,一份临床记录可能包含30-40个描述患者病情的数据条目,但一个完整的基因组序列可能揭示10,000个癌症突变。如果简单地结合起来,突变的绝对数量可能会压倒临床信息的重要性,这可能会导致分析和解释的偏差,例如,未能考虑重要的社会经济或种族信息。我们将制定方法,使不同数据来源的重要性相等,以便它们可以以公平和公正的方式结合起来。该项目福尔斯属于EPSRC医疗保健技术研究领域,其中“优化疾病预测,诊断和干预”是本网站上列出的主题或研究领域之一。它将创建分析大型数据集的新方法,支持患者特定的预测模型,并支持识别预防疾病或其复发的机会。该项目将涉及与牛津大学的合作-癌症免疫学公司Singula Bio。
英文摘要
Ovarian cancer is the 6th most common cancer for women in the UK. High-grade serous ovarian cancer (HGSOC) accounts for most cases, with a low 30% 5-year survival. The two main factors that contribute to this poor prognosis are: 1) late diagnosis of the disease, and 2) a high proportion of relapse despite initial response to treatment. The latter suggests that small populations of treatment resistant cancer cells may exist that can repopulate the disease. It is therefore of interest to identify such cancer cells and understanding how they different from other cancer cell types that might also affect why patients respond differently to treatment and differ in how long they survive. The molecular basis of ovarian cancer can be unravelled using a plethora of modern technologies such as sequencing and imaging at both bulk tissue and single-cell level. This is creating an unprecedented opportunity to use a data-driven approach to enable the precise characterisation of ovarian cancer and the possibility of developing targeted treatment options. However, to make effective use of the molecular data, robust analytical approaches are required to characterise cell populations of interest (particularly rare ones) and to integrate heterogeneous data modalities.This research aims to:1) To develop a robust and interpretable approach to identify rare cell populations from high-dimensional molecular data,2) To develop a statistical framework for the integration of multimodal data for survival prediction.Novelty of the research methodology We will develop statistical techniques that are specifically designed to identify rare cell types from molecular data. Classical statistical discovery methods tend to be biased toward the most common cell populations as there is more information about them. There is often a penalty associated with suggesting a rare cell type as these may not be real so a balance must be struck between proposing new cell types and the chance that these proposals may turn out to be false after further examination. We will develop techniques that allow us to control the balance between these competing needs allowing experimental scientists to adjust expectations based on the level of acceptable risk available to them. We will also develop techniques to combine different sources of data such as clinical record, magnetic resonance imaging and whole genome sequencing. These techniques will examine a specific limitation of existing approaches which typically do not account for the information imbalance between different types of data. For example, a clinical record might contain 30-40 data entries describing a patient's condition, but a whole genome sequence might reveal 10,000s of cancer mutations. If naively combined, the sheer number of mutations can overwhelm the importance of the clinical information, which can cause biases in analysis and interpretation, for instance, by failing to consider important socioeconomic or ethnicity information. We will develop approaches that equalise the important placed on different sources of data such that they can be combined in a fair and equitable way. This project falls within the EPSRC Healthcare Technologies research area' where "Optimising disease prediction, diagnosis and intervention" is one of the themes or research areas listed on this website.It will create new methods for analysing large data sets, underpin patient-specific predictive models, and support the identification of opportunities for prevention of disease or its recurrence.This project will involve a collaboration with the Oxford-based cancer immunology company, Singula Bio.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于随机网络演算的无线机会调度算法研究
  • 批准号:
    60702009
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2007
  • 负责人:
    雷蕾
  • 依托单位: