课题基金 / 基金详情

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该建议将开发一些新的统计工具,用于从实验数据中学习基因型-表型映射。大量的基因型-表型数据集可以通过遗传多样化产生,然后通过高通量筛选/选择和功能不同群体的下一代DNA测序。由此产生的数据提出了新的和有趣的统计挑战,包括大量的例子,只存在的反应,和噪声/缺失的数据。仅存在响应的出现是因为大多数高通量筛选/选择方法仅分离功能性实例(阳性响应),而非功能性实例(阴性)难以或不可能获得。所得数据集包含初始未标记的变体文库和阳性实例。本提案中开发的建模工具适用于从分子到生态系统的所有生物组织水平。该提案中开发的新统计方法将模拟蛋白质序列,结构和功能之间的关系,目的是深入了解生化机制并设计新的有用蛋白质。 该提案将(i)开发新的理论和工具来分析由新兴高通量方法生成的大量蛋白质序列功能数据;(ii)解决与阳性未标记(PU)学习,超大数据量,低质量/缺失数据相关的挑战,以及 (iii)编码来自现有数据库或物理模型的辅助信息。此外,应用这项工作中开发的方法和算法将产生新的科学见解和工程生物系统。
英文摘要
This proposal will develop a number of novel statistical tools for learning genotype-phenotype mappings from experimental data. Massive genotype-phenotype data sets can be generated by genetic diversification, followed by high-throughput screening/selection and next-generation DNA sequencing of functionally-distinct populations. The resulting data presents new and interesting statistical challenges including large numbers of examples, presence-only responses, and noisy/missing data. Presence-only responses arise because most high-throughput screening/selection methods isolate only functional examples (positive responses), while non-functional examples (negatives) are difficult or impossible to obtain. The resulting data sets contain the initial unlabelled variant library and positive examples. The modeling tools developed in this proposal apply to all levels of biological organization spanning from molecules to ecosystems. The novel statistical methods developed in this proposal will model the relationships between protein sequence, structure, and function, with the goal of gaining insight into biochemical mechanisms and designing new and useful proteins. This proposal will (i) develop new theory and tools to analyze the large quantities of protein sequence­ function data that are being generated by emerging high-throughput methods; (ii) address challenges associated with positive-unlabeled (PU) learning, extremely large data size, low- quality/missing data, and (iii) encoding side information from existing databases or physical models. Furthermore, applying the methods and algorithms developed in this work will generate novel scientific insights and engineered biological systems.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
PUlasso: High-Dimensional Variable Selection With Presence-Only Data
PUlasso:仅存在数据的高维变量选择
DOI: 10.1080/01621459.2018.1546587
发表时间: 2018
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Song, Hyebin, Raskutti, Garvesh]
通讯作者: Raskutti, Garvesh
DOI: 10.1214/20-ejs1677
发表时间: 2020-01-01
期刊: ELECTRONIC JOURNAL OF STATISTICS
影响因子: 1.1
作者: [Dai, Ran, Song, Hyebin, Raskutti, Garvesh]
通讯作者: Raskutti, Garvesh
海外基金