BSF:2012304:Methods for Preprocessing Population Sequence Data
BSF:2012304:Methods for Preprocessing Population Sequence Data
批准号:
1331176
负责人:
Eleazar Eskin
金额:
$4.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-09-01 至 2018-08-31
中文摘要
该项目是美国-以色列计算机科学合作(USICCS)计划的一部分。通过这一计划,NSF和美国-以色列双国科学基金会(BSF)共同支持美国研究人员和以色列研究人员之间的合作。近年来,进行了许多基因研究,揭示了人类基因变异与复杂疾病之间的许多新联系。这些研究被称为全基因组关联研究,仅限于常见的遗传变异,因为收集遗传变异的技术仅限于收集常见的变异。有证据表明,罕见的变异在疾病结构中发挥着重要作用。最近,测序技术已经被引入,它能够收集遗传常见和罕见的遗传变异。测序技术产生了海量数据,带来了新的计算挑战。在这个项目中,个人信息系统将开发解决这些计算挑战的方法,包括设计有效的算法和对排序过程进行建模。此外,研究人员将开发将稀有变异纳入遗传研究分析的方法。我们项目的直接更广泛的影响是这些工具可供遗传学家普遍使用,导致对疾病遗传学的更好理解。特别是,PI将把他们的方法应用于非霍奇金淋巴瘤、躁郁症、血脂异常、神经退行性痴呆和图雷特综合征的研究,这将直接影响我们对这些特殊情况的理解。目前用于分析测序数据的计算方法是存在的,但它们仅限于对单个样本的分析。在这个项目中,PI将设计有效的计算方法来分析整个种群的序列数据。对于总体样本,巨大的数据量要求在内存和运行时间方面设计高效的算法。具体地说,PI建议设计用于压缩测序数据的算法,用于搜索在多个样本中通过下降而相同的区域,以及用于从序列数据进行高分辨率单倍型推断。PI将显式地对稀有变异和测序过程进行建模,并使用机器学习技术和凸优化来有效地估计模型参数。这些方法将允许对人口数据进行精细的分析,从而改善对复杂疾病和人类历史的理解。该项目的合作性质将使参与该项目的学生接触到以色列和美国的医学和遗传学世界,并将提高他们设计和实施复杂算法问题解决方案的能力。在这个项目中开发的方法将成为加州大学洛杉矶分校和特拉维夫大学课程教材的一部分,这些教材将公开提供。
英文摘要
This project is funded as part of the United States-Israel Collaboration in Computer Science (USICCS) program. Through this program, NSF and the United States - Israel Binational Science Foundation (BSF) jointly support collaborations among US-based researchers and Israel-based researchers.In recent years, many genetic studies have been performed, revealing many new associations between human genetic variation and complex diseases. These studies, referred to as genome-wide association studies, are limited to common genetic variants because the technology which collected the genetic variation was limited to only collecting common variants. There is evidence suggesting that rare variants have an important role in disease architectures. Recently, sequencing technologies have been introduced which are capable of collecting both genetic common and rare genetic variation. Sequencing technologies generate enormous amounts of data, raising new computational challenges. In this project, the PIs will develop methods for addressing these computational challenges including the design of efficient algorithms and the modeling of the sequencing process. In addition, the researchers will develop methods for incorporating rare variants into the analysis of genetic studies. The immediate broader impact of our project is the availability of these tools for general use by geneticists, leading to an improved understanding of the disease genetics. Particularly, the PIs will apply their methods to studies of non-Hodgkin's lymphoma, bipolar, dyslipidemia, neurodegenerative dementia, and Tourette syndrome, which will result in a direct impact on our understanding of these particular conditions.Current computational methods for the analysis of sequencing data exist, however they are limited to the analysis of a single sample. In this project the PIs will design efficient computational methods for the analysis of sequence data across a population. For population samples, the tremendous size of the data requires the design of highly efficient algorithms in terms of memory and runtime. Specifically, the PIs propose to design algorithms for the compression of sequencing data, for the search of regions identical by descent across multiple samples, and for high-resolution haplotype inference from sequence data. The PIs will explicitly model rare variants and the sequencing process, and use machine learning techniques and convex optimization to estimate the model parameters efficiently. These methods will allow for a fine-scale analysis of population data, resulting in improved understanding of complex diseases and human history. The collaborative nature of the project will expose the students involved in the project to the medical and genetics worlds, both in Israel and in the US, and it will improve their abilities to design and implement solutions to complex algorithmic problems. The methods developed in this project will be part of the teaching material of courses in UCLA and Tel-Aviv, and these materials will be made publicly available.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Medium: Causal inference in biobanks: Leveraging genetics to infer causal relationships using electronic health records
-
批准号:2106908
-
项目类别:Continuing Grant
-
资助金额:$119.99万
-
财政年份:2021
-
负责人:Eleazar Eskin
-
依托单位:
III:Small: Replication Studies for High Dimensional Data: Insights into Confounding and Heterogeneity
-
批准号:1910885
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2019
-
负责人:Eleazar Eskin
-
依托单位:
III: Medium: Detecting Low Dimensional Structures in Genomic Data
-
批准号:1705197
-
项目类别:Standard Grant
-
资助金额:$119.97万
-
财政年份:2017
-
负责人:Eleazar Eskin
-
依托单位:
III: Small: Causal and Statistical Inference in the Presence of Confounding Factors
-
批准号:1320589
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2013
-
负责人:Eleazar Eskin
-
依托单位:
III: Medium: Meta-analysis reinterpreted using causal graphs
-
批准号:1302448
-
项目类别:Continuing Grant
-
资助金额:$112.08万
-
财政年份:2013
-
负责人:Eleazar Eskin
-
依托单位:
III: Medium: Private Identification of Relatives and Private GWAS: First Steps in the New Field of CryptoGenomics
-
批准号:1065276
-
项目类别:Standard Grant
-
资助金额:$70.0万
-
财政年份:2011
-
负责人:Eleazar Eskin
-
依托单位:
III: Small: Inference of Causal Regulatory Relationships from Genetic Studies
-
批准号:0916676
-
项目类别:Continuing Grant
-
资助金额:$49.94万
-
财政年份:2009
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: Design and Analysis of Compressed Sensing DNA Microarrays
-
批准号:0729049
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2007
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: SEIII: Estimating Haplotype Frequencies
-
批准号:0731455
-
项目类别:Standard Grant
-
资助金额:$13.59万
-
财政年份:2007
-
负责人:Eleazar Eskin
-
依托单位:
Collaborative Research: SEIII: Estimating Haplotype Frequencies
-
批准号:0513612
-
项目类别:Standard Grant
-
资助金额:$29.5万
-
财政年份:2005
-
负责人:Eleazar Eskin
-
依托单位: