Physicochemical models of protein-DNA binding with standard and modified base pairs.
Physicochemical models of protein-DNA binding with standard and modified base pairs.
复制标题
蛋白质-DNA 与标准碱基对和修饰碱基对结合的物理化学模型。
DOI:
10.1073/pnas.2205796120
复制
发表时间:
2023-01-24
影响因子:
11.1
通讯作者:
中科院分区:
文献类型:
--
作者:
The Watson–Crick sequence model enables a simplified representation of DNA, wherein four letters, A, C, G, and T, describe the chemical identities and orientations of all possible nucleotide pairs. In this coarse-grained model, each letter describes an assembly of over 60 atoms. However, the atomic composition of a nucleotide pair can be altered by chemical modifications, different base-pairing geometries, or mismatches. As only a few atoms contribute to binding specificity, we propose that compared to a sequence model, a chemistry-based model that directly encodes protein–DNA contacts may more robustly capture the chemical variations of DNA. We introduce models that directly and precisely represent physicochemical readout, which is, importantly, not restricted to standard Watson–Crick base pairs. DNA-binding proteins play important roles in various cellular processes, but the mechanisms by which proteins recognize genomic target sites remain incompletely understood. Functional groups at the edges of the base pairs (bp) exposed in the DNA grooves represent physicochemical signatures. As these signatures enable proteins to form specific contacts between protein residues and bp, their study can provide mechanistic insights into protein–DNA binding. Existing experimental methods, such as X-ray crystallography, can reveal such mechanisms based on physicochemical interactions between proteins and their DNA target sites. However, the low throughput of structural biology methods limits mechanistic insights for selection of many genomic sites. High-throughput binding assays enable prediction of potential target sites by determining relative binding affinities of a protein to massive numbers of DNA sequences. Many currently available computational methods are based on the sequence of standard Watson–Crick bp. They assume that the contribution of overall binding affinity is independent for each base pair, or alternatively include dinucleotides or short k-mers. These methods cannot directly expand to physicochemical contacts, and they are not suitable to apply to DNA modifications or non-Watson–Crick bp. These variations include DNA methylation, and synthetic or mismatched bp. The proposed method, DeepRec, can predict relative binding affinities as function of physicochemical signatures and the effect of DNA methylation or other chemical modifications on binding. Sequence-based modeling methods are in comparison a coarse-grain description and cannot achieve such insights. Our chemistry-based modeling framework provides a path towards understanding genome function at a mechanistic level.
登录
查看更多内容
影响因子:
46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者:
Bulyk, Martha L.
影响因子:
48
作者:
Isakova, Alina;Groux, Romain;Deplancke, Bart
通讯作者:
Deplancke, Bart
影响因子:
56.9
作者:
AGGARWAL, AK;RODGERS, DW;HARRISON, SC
通讯作者:
HARRISON, SC
影响因子:
13.8
作者:
Liu, Yiwei;Zhang, Xing;Blumenthal, Robert M.;Cheng, Xiaodong
通讯作者:
Cheng, Xiaodong
影响因子:
3.9
作者:
Chatterjee, Raghunath;He, Ximiao;Vinson, Charles
通讯作者:
Vinson, Charles