CAREER: Optimizing Learning Models for Interpretation of Heterogeneous Biological Data
CAREER: Optimizing Learning Models for Interpretation of Heterogeneous Biological Data
批准号:
1453658
负责人:
Tomasz Arodz
金额:
$43.57万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-02-01 至 2021-01-31
中文摘要
快速收集数据的技术与浏览由此产生的大型异构数据集并制定新假设的能力之间的差距正在扩大。在分子生物学中,这限制了基础发现,并减缓了结果从实验室到临床的转化。弥合这一差距必然需要智能算法。然而,尽管取得了重大进展,机器学习算法仍难以承担支撑有希望的假设的计算风险。一台机器可以重新生成许多与实验数据相匹配的假设,但考虑到我们对生物学的理解,它们可信吗?一种算法可以利用这些数据对疾病中哪些已知的生物通路失调进行评分,但它能对观察结果提出一种新的解释吗?该项目将通过新颖的算法解决当前方法的局限性,这些算法将能够将现有知识集成到数据驱动的分析和建模中。在涉及真实和合成数据集的严格验证之后,所提出的算法将为生物学家形成一个变革性的计算工具集,极大地增强了理解涉及异质元素以复杂方式相互作用的分子系统的正常和病理行为的能力。算法方面的工作将涉及本科生和研究生。通过参与多学科研究,他们将学习如何克服跨学科交流中的障碍,这在计算机方法渗透到许多其他领域的世界中至关重要。学生还将受益于图论和机器学习的新课程。PI将把课程中的学习模块提供给其他机构的教师。此外,每年,PI将组织一场高中编程竞赛,作为算法如何帮助解决社会和科学问题的实际演示。高中、本科和研究生阶段的所有活动都将侧重于支持妇女和其他代表性不足的群体探索计算机科学。该项目的主要目标是创建训练分类器的算法,这些算法是准确和可解释的,将通过将集成、子模块集函数和非光滑函数优化技术联系在一起的图正则化机器学习来实现。近来,次模正则器作为一种为线性模型提供特征空间结构信息的有效方法受到了人们的关注,但将这种方法扩展到非线性分类器中受到了测量特征与预测结果之间关系的非线性的阻碍。本研究将通过设计创新的基于集成的方法来克服这一障碍,该方法结合了子模块正则化器,从而提高了准确性,减少了过度训练,增加了模型的可解释性。这些新方法将适用于任何涉及由网络连接的特征的分类问题。这样的问题并不局限于生物学;这些方法将有更广泛的用途,例如图像分析或文本分类。对于生物学应用,这将是工作的主要焦点,将构建一个统一的元网络,将遗传,表观遗传,转录组学,蛋白质组学和代谢组学水平的分子元件联系起来,并将使用谱图理论的概念设计一种处理缺失数据的新方法。
英文摘要
The gap between techniques for rapid gathering of data and the ability to navigate the resulting large, heterogeneous datasets and formulate novel hypotheses is widening. In molecular biology, this limits basic discoveries and slows the translation of results from laboratory to clinic. Bridging this gap will necessarily involve intelligent algorithms. Yet, despite significant advances, machine-learning algorithms struggle with taking the calculated risk that underpins promising hypotheses. A machine can generate de novo a number of hypotheses that match experimental data, but are they plausible given our understanding of biology? An algorithm can use the data to score which of the known biological pathways is dysregulated in the disease, but can it propose a novel explanation of the observations? This project will address the limitation of current methods through novel algorithms that will be able to integrate existing knowledge into data-driven analysis and modeling. Following rigorous validation involving real and synthetic datasets, the proposed algorithms will form a transformative computational tool set for biologists, greatly enhancing the capabilities for understanding normal and pathological behavior of molecular systems involving heterogeneous elements interacting in complex ways. The work on the algorithms will involve undergraduate and graduate students. Through participation in multidisciplinary research, they will learn how to overcome barriers in cross-discipline communication, which is crucial in the world where computer methods permeate many other areas. Students will also benefit from a new course on Graph Theory and Machine Learning. The PI will make the learning modules from the course available online to instructors at other institutions. Also, annually, the PI will organize a high-school programming contest that will serve as a hands-on demonstration of how algorithms can help solve societal and scientific problems. All activities at high-school, undergraduate and graduate levels will have emphasis on supporting women and other underrepresented groups in their exploration of computer science.The project's main objective of creating algorithms for training classifiers that are accurate and interpretable will be achieved through graph-regularized machine learning that ties together ensembles, submodular set functions, and techniques for non-smooth function optimization. Submodular regularizers have recently received attention as a powerful way for equipping linear models with information about structures in the feature space, but extending the approach to non-linear classifiers is hampered by the nonlinearity of the relationship between measured features and the predicted outcome. This research will move past this obstacle by designing innovative ensemble-based methods that incorporate submodular regularizers, thus leading to improved accuracy, reduced overtraining and increased interpretability of models. These novel methods will be applicable to any classification problem that involves features connected by a network. Such problems are not confined to biology; the methods will be of broader use, for example to image analysis or text categorization. For biological applications, which will be the main focus of the work, a unified meta-network that links molecular elements at genetic, epigenetic, transcriptomic, proteomic and metabolomics levels will be constructed, and a new approach for dealing with missing data will be designed using concepts from spectral graph theory.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
AF: Small: Approximation algorithms for quantum mechanical problems
-
批准号:1617710
-
项目类别:Standard Grant
-
资助金额:$38.08万
-
财政年份:2016
-
负责人:Tomasz Arodz
-
依托单位:
海外基金