Statistical Tools for Post-Genomic Personalized Medicine and Health Care
Statistical Tools for Post-Genomic Personalized Medicine and Health Care
批准号:
1001643
负责人:
Mathukumalli Vidyasagar
金额:
$30.45万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-01 至 2015-05-31
中文摘要
这项拟议研究的目标是为系统生物学开发一些统计工具,以加速后基因组时代个性化医疗和保健的发展。将研究药物发现过程中特别适合系统生物学方法的那些方面,即:靶标和铅的识别,以及临床试验的患者选择。相关的统计问题是:检测序列中的相似性,以及在保持其统计特性的同时对大量数据进行二次抽样。为了检测序列相似性,提出将广泛使用的BLAST(基本局部比对搜索技术)算法扩展到被比较的序列是马尔可夫链的样本路径的情况,现有的BLAST理论假设它们是I.I.D.的样本路径。流程。为了在保持统计特性的同时进行二次抽样,建议使用Vapnik-Chervonenkis(VC-)理论和VC-维。在未来的某个日期,这个项目之后将是一个更雄心勃勃的多PI项目,它将把这里开发的方法应用到特定的治疗领域。技术价值个人医学是系统生物学研究的一个有价值的目标。在个性化医学中,确定合适的药物靶点和候选药物(称为先导药物)的问题非常适合于系统生物学方法,就像基因-表型相关性(将一个人的基因变异与他/她的疾病倾向、药物反应等相关)一样。因此,这里选择的问题是药物发现中最自然的问题,需要通过系统生物学的方法来解决。目标和前导的识别是通过序列比较进行的,因此寻找考虑符号之间可能的马尔可夫相互依赖的方法是自然的。目前广泛使用的BLAST算法是基于大偏差理论的一种改进,因此提出利用PI最近对马尔可夫链的大偏差理论的简化扩展来获得BLAST理论在马尔可夫依赖样本路径上的扩展。Vapnik-Chervonenkis(VC-)理论是唯一“无分布”的抽样理论;也就是说,保持原始数据的统计特性所需的子样本的大小不取决于原始数据的分布。因此,这一理论非常适合于次抽样问题。事实上,已经取得了一些初步结果。广泛的影响拟议的研究将对其他科学领域产生影响,如互联网计算、说话人识别等,因为通过次抽样检测序列相似性和减少数据量的问题出现在各种背景下,而不仅仅是在系统生物学中。为了促进教育外展,国际和平协会将继续他过去通过广泛使用的教科书和专著传播知识的做法。已经与普林斯顿大学出版社签署了一份合同,将在研究生水平的专著中发表拟议的研究结果。一旦理论研究完成,BLAST对马尔可夫链的扩展将在一个用户友好的软件包中实现,该软件包具有与原始BLAST相同的图形用户界面,并且可以免费下载。扩展的BLAST实施还将托管在德克萨斯大学达拉斯分校的生物计算集群上。PI将继续向一年级本科生教授计算生物学入门课程,就像他过去所做的那样。PI将参与UTD的UTeach计划,以编写高中水平的计算生物学教材。他还将通过访问当地学校和谈论他的研究的社会意义来积极招募妇女和少数族裔。
英文摘要
The objective of the proposed research is to develop some statistical tools for systems biology that will accelerate the development of personalized medicine and health care in the post-genomic era. Those aspects of the drug discovery process that are particularly suited to a systems biologyapproach will be studied, namely: target and lead identication, and patient selection for clinical trials. The associated statistical problems are: detecting similarity in sequences, and subsampling enormous-sized data while retaining its statistical properties. For detecting sequence similarity, itis proposed to extend the widely-used BLAST (Basic Local Alignment Search Technique) algorithm to the case where the sequences being compared are sample paths of Markov chains; the existing BLAST theory assumes that they are sample paths of i.i.d. processes. For subsampling while retaining statistical properties, it is proposed to use Vapnik-Chervonenkis (VC-) theory and VC-dimension. At some future date, this project will be followed by a more ambitious multi-PI project that will apply the methods developed here to specific therapeutic areas.Technical MeritPersonal medicine is a worthy goal for systems biology research. Within personalized medicine, the problems of identifying suitable targets for drugs, and candidate drugs (known as leads) are quite amenable to a systems biology approach, as is the problem of genotype-phenotype correla-tion (correlating a person's genetic variation with his/her propensity to disease, responsiveness to drugs, etc.). Thus the problems chosen here are the most natural problems in drug discovery to be addressed via a systems biology approach. Target and lead identification proceeds via sequence comparison, so it is natural to seek methods that take into account the possible Markovian interdependence amongst the symbols. The widely-used BLAST algorithm is based on a modification oflarge deviation theory, so it is proposed to use the PI's recent simplified extension of large deviation theory for Markov chains to obtain an extension of BLAST theory to Markov-dependent sample paths. Vapnik-Chervonenkis (VC-) theory is the only theory of sampling that is `distribution-free' ;that is, the size of the subsample needed to retain the statistical properties of the original data does not depend on the distribution of the original data. Hence this theory is ideally suited to the problem of subsampling. Indeed some preliminary results have already been obtained.Broader ImpactThe proposed research will have impact on other areas of science such as Internet computing,speaker recognition etc., because problems of detecting sequence similarity and reducing data sizes by subsampling arise in a variety of contexts, and not just in systems biology. To promote educationoutreach, the PI will continue his past practice of disseminating knowledge through widely-used textbooks and monographs. A contract has already been signed with Princeton University Press to publish the outcomes of the proposed research in a graduate level monograph. Once the theoreticalresearch is completed, the extension of BLAST to Markov chains will be implemented in a user-friendly software package that has the same GUI as the original BLAST, and be made freely downloadable. The extended BLAST implementation will also be hosted on a bio-computing cluster at UT Dallas. The PI will continue to teach introductory computational biology to first-year undergraduates as he has done in the past. The PI will participate in the UTeach program of UTD to generate educational materials for computational biology at the High School level. He will also actively recruit women and minorities by visiting area schools and speaking about the socialrelevance of his research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms and Statistical Methods for Personalized Diagnosis and Therapy in Cancer
-
批准号:1306630
-
项目类别:Standard Grant
-
资助金额:$36.96万
-
财政年份:2013
-
负责人:Mathukumalli Vidyasagar
-
依托单位:
海外基金