Algorithms and Statistical Methods for Personalized Diagnosis and Therapy in Cancer
Algorithms and Statistical Methods for Personalized Diagnosis and Therapy in Cancer
批准号:
1306630
负责人:
Mathukumalli Vidyasagar
金额:
$36.96万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-07-01 至 2018-06-30
中文摘要
目前,有几项公共努力正在进行,以生成来自所有可用癌症组织的大量数据集。卵巢癌、肺癌、乳腺癌和结肠癌等四种类型的癌症已经有了这样的数据,更多的癌症正在研究中。这些数据集的一个特点是,测量的特征数量是数万个,而每种癌症的组织样本数量是数百个。因此,主要的挑战是提取出最具信息量的特征,这些特征可以用来区分一组癌症患者和另一组癌症患者,例如,区分那些对特定形式的治疗有反应的患者和那些没有反应的患者。这些特征被称为生物标志物,可以用于开发针对特定群体甚至个体患者的定制疗法。然而,几乎可用的从大数据集中提取相关特征的算法都面临着一个“障碍”,即提取的特征数量受到训练样本数量的限制。这个数字可能有数百个,太大了,无法在生物学应用中发挥作用。在这个项目中,我们提出了一些新的特征提取算法,可以突破这个“障碍”,识别远少于训练样本数量的特征。这些新开发的算法将从其统计行为和最优性方面进行分析;此外,它们还将在肺癌、卵巢癌和子宫内膜癌的实际数据集上进行验证。当前癌症治疗的另一个重要方面是广泛接受使用多种药物联合治疗的需要。这是因为当患者只用一种药物治疗时,即使肿瘤最初缩小,几乎总是会重新长出来,而且复发的肿瘤通常对药物有抗药性。由于组合爆炸,在实验环境中尝试所有可能的药物组合是不可行的。此外,由于癌细胞行为的复杂性,也不可能建立多种药物联合使用的作用机制的分析模型。因此,在对每种药物的作用机制几乎不做任何假设的情况下,开发预测多种药物联合疗效的方法是必要的。在本项目中,建议使用所谓的“最大熵法”来开发这样的预测方法。最大熵法是在信息论中导出均衡统计机制的背景下发展了大约50年,并且被广泛接受为当需要最小化先验假设数量时使用的最佳方法之一。智力优势:目前可用的分类和回归算法,如LASSO、elasticnet和Dantzig,其特征是提取的关键特征数量大致等于训练样本的数量。然而,即使这个数字太大,在生物学情况下也无法实际使用。对PI发明的一种新算法的初步研究表明,它不存在这种限制。此外,该算法在子宫内膜癌和卵巢癌两种类型的癌症数据集上显示出良好的性能。如果可以为该算法的观察行为以及另一个仍处于概念阶段的算法建立一个良好的理论基础,这将是对统计学和机器学习理论的一个非常重要的贡献。另一方面,如果能够通过理论和实验建立最大熵法预测多药联合疗效的方法,将大大提高该方法的理论水平和多药联合治疗的实用性。更广泛的影响:癌症是美国、其他工业化国家和新兴工业化国家的第二大死因。人们普遍认为癌症是最“独特”的疾病,因为没有两种疾病的表现是相同的。因此,个性化治疗是前进的方向。然而,很少有方法发展个人治疗是不可知论癌症的类型。本项目旨在发展这种方法。考虑到癌症在科学界和整个社会中所占的很大比重,可以肯定的是,如果该项目成功完成,那么癌症研究人员将会对研究结果进行跟进。为了加快这一进程,PI将与德克萨斯大学西南医学中心和休斯顿安德森癌症中心的几位癌症研究人员合作。该项目每年将培训两名研究生和至少一名本科生暑期实习生。这将有助于增加训练有素的人力资源,并将癌症治疗设计的分析方法传播给更广泛的受众。
英文摘要
At present, there are several public efforts under way to generate massive data setsderived from all available cancer tissues. Such data is already available for four forms of cancer: Ovarian,lung, breast and colon, and more are on the way. One of the characteristics of these data sets is that thenumber of features that are measured is in the tens of thousands, while the number of tissue samples foreach form of cancer is in the hundreds. The main challenge therefore is to extract the most informative featuresthat can be used to distinguish one set of cancer patients from another, for example, those that respond to aparticular form of therapy from those who do not. Such features, referred to as biomarkers, can then be usedto develop therapies that are customized to focused groups or even individual patients. However, almostall available algorithms for extracting relevant features from big data sets face a "barrier" in that the numberof features extracted is bounded below by the number of training samples. This number, which might bein the hundreds, is far too large to be useful in biological applications. In this project, it is proposed todevelop some novel algorithms for feature extraction that can break through this "barrier" and identify farfewer features than the number of training samples. These newly developed algorithms will be analyzedin terms of their statistical behavior and their optimality; in addition they will be validated on actual datasets from lung, ovarian and endometrial cancer.Another important aspect of current cancer therapy is the widespread acceptance of the need to usemulti-drug combinations. This is because when a patient is treated with a single drug, almost invariablythe tumor will grow back even if it shrinks initially, and the relapsed tumor is often resistant to the drug.Due to combinatorial explosion, it is not feasible to try out all possible combinations of drugs in experimentalsettings. Moreover, due to the complexity of the behavior of cancer cells, it is also not possibleto develop analytical models for the mechanisms of action of multiple drugs used in combination. It istherefore imperative to develop methodologies for predicting the efficacy of multi-drug combinations whilemaking almost no assumptions about the mechanism of action of each drug. In this project, it is proposed to usethe so-called "maximum entropy method" to develop such a prediction methodology. The maximum entropymethod was developed about fifty years in the context of deriving equilibrium statistical mechanicsfrom information theory, and is widely accepted as one of the best methods to be used when it is desired tominimize the number of a priori assumptions.Intellectual Merit: Currently available algorithms for classification and regression such as LASSO, elasticnet, and Dantzig have the feature that the number of key features extracted is roughly equal to the numberof training samples. However, even this number is too large to be of practical use in biological situations.Preliminary investigations on a new algorithm invented by the PI show that it does not have this limitation.Moreover, this new algorithm has shown promising performance on two types of cancer data sets:endometrial and ovarian. If a sound theoretical foundation can be established for the observed behavior ofthis algorithm, as well as for another that is still in the conceptual stage, that would be a very significantcontribution to statistics and to machine learning theory. On another front, if it can be established throughtheory and experiment that the maximum entropy method can be used to predict the efficacy of multi-drugcombinations, that would greatly advance both the theory of the method and the practical applicability ofmulti-drug therapy.Broader Impacts: Cancer is the second leading cause of death in the USA, in other industrialized countries,and also in newly industrializing countries. It is widely accepted that cancer is the most "individual" of diseasesin that no two manifestations are alike. Therefore personalized therapy is the way forward. However,there are very few methodologies for developing personal therapy that are agnostic as to the type of cancer.The present project aims to develop precisely such methodologies. Given the large mindshare of cancer inthe scientific community and in society at large, it can be safely assumed that if the project is successfullycompleted, then the research findings would be followed up by the cancer researcher community. To hastenthe process, the PI will work with several cancer researchers in the UT Southwestern Medical Center inDallas and in the M. D. Anderson Cancer Center in Houston.The project will entail the training of two graduate students and at least one undergraduate summerintern per year. This would serve to increase the pool of trained manpower and also to disseminate theanalytical approach to cancer therapy design to a broader audience.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Tools for Post-Genomic Personalized Medicine and Health Care
-
批准号:1001643
-
项目类别:Standard Grant
-
资助金额:$30.45万
-
财政年份:2010
-
负责人:Mathukumalli Vidyasagar
-
依托单位:
海外基金