Inferring Protein Sequence-Function Relationships with Large-Scale Positive-Unlabeled Learning.

Inferring Protein Sequence-Function Relationships with Large-Scale Positive-Unlabeled Learning.
复制标题

推断蛋白质序列功能与大规模阳性未标记的学习关系。

DOI:
10.1016/j.cels.2020.10.007
复制
发表时间:
2021-01-20
期刊:
影响因子:
9.3
通讯作者:
Romero PA
Romero PA
中科院分区:
生物学1区
文献类型:
--
作者:
Song H;Bremer BJ;Hinds EC;Raskutti G;Romero PA

文献摘要

参考文献

相似文献

机器学习可以推断蛋白质序列如何映射到功能,而不需要详细了解潜在的物理或生物机制。将现有的监督学习框架应用于深度突变扫描(DMS)和相关方法生成的大规模实验数据具有挑战性。 DMS 数据通常包含高维和相关序列变量、实验采样误差和偏差以及缺失数据的存在。值得注意的是,大多数 DMS 数据不包含负序列的示例,这使得直接估计序列如何影响功能具有挑战性。在这里,我们开发了一个正向无标记 (PU) 学习框架,用于从大规模 DMS 数据中推断序列函数关系。我们的 PU 学习方法在十个大规模序列功能数据集(代表不同折叠、功能和库类型的蛋白质)中显示出出色的预测性能。估计的参数精确定位了决定蛋白质结构和功能的关键残基。最后,我们应用统计序列功能模型来设计高度稳定的酶。随着高通量实验的进步,蛋白质序列功能数据的数量正在迅速增长。宋等人。提出了一种机器学习方法,可以从深度突变扫描生成的大规模数据中推断序列-功能关系。学习到的模型捕获了蛋白质结构和功能的重要方面,可用于设计新的和增强的蛋白质。
Machine learning can infer how protein sequence maps to function without requiring a detailed understanding of the underlying physical or biological mechanisms. It’s challenging to apply existing supervised learning frameworks to large-scale experimental data generated by deep mutational scanning (DMS) and related methods. DMS data often contain high dimensional and correlated sequence variables, experimental sampling error and bias, and the presence of missing data. Notably, most DMS data do not contain examples of negative sequences, making it challenging to directly estimate how sequence affects function. Here, we develop a positive-unlabeled (PU) learning framework to infer sequence-function relationships from large-scale DMS data. Our PU learning method displays excellent predictive performance across ten large-scale sequence-function data sets, representing proteins of different folds, functions, and library types. The estimated parameters pinpoint key residues that dictate protein structure and function. Finally, we apply our statistical sequence-function model to design highly stabilized enzymes. The quantity of protein sequence-function data is growing rapidly with advances in high-throughput experimentation. Song et al. present a machine learning approach to infer sequence-function relationships from large-scale data generated by deep mutational scanning. The learned models capture important aspects of protein structure and function, and can be applied to design new and enhanced proteins.
DOI: 10.1038/s41586-018-0461-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Findlay GM;Daza RM;Martin B;Zhang MD;Leith AP;Gasperini M;Janizek JD;Huang X;Starita LM;Shendure J
通讯作者: Shendure J
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
使用机器学习和合成基因进行工程蛋白酶K。
DOI: 10.1186/1472-6750-7-16
发表时间: 2007-03-26
期刊: BMC BIOTECHNOLOGY
影响因子: 3.5
作者:
Liao, Jun;Warmuth, Manfred K.;Govindarajan, Sridhar;Ness, Jon E.;Wang, Rebecca P.;Gustafsson, Claes;Minshull, Jeremy
通讯作者: Minshull, Jeremy
DOI: 10.1038/nmeth.3027
发表时间: 2014-08
期刊: NATURE METHODS
影响因子: 48
作者:
Fowler, Douglas M.;Fields, Stanley
通讯作者: Fields, Stanley
DOI: 10.2307/1390605
发表时间: 2000-03-01
影响因子: 2.4
作者:
Lange, K;Hunter, DR;Yang, I
通讯作者: Yang, I