Method Development: Efficient Computer Vision Based Algorithms
Method Development: Efficient Computer Vision Based Algorithms
批准号:
7965320
负责人:
Ruth Nussinov
金额:
$13.03万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AddressAlgorithmsArtsBindingBinding SitesBiologicalCaliberCatalysisCatalytic DomainCellsCharacteristicsChemical EngineeringClassificationCollectionComplementComputational BiologyComputer SimulationComputer Vision SystemsCore ProteinCytochrome P450Data SetDatabasesDetectionDevelopmentDimensionsDockingElectrolytesEnzymesGoalsHumanIntegral Membrane ProteinInternetIon ChannelLibrariesLifeLigandsLinkLocationMapsMedialMembraneMembrane ProteinsMethodologyMethodsModelingMolecularMonitorMovementNatureOperative Surgical ProceduresOrganismOutputPathway interactionsPerformancePhage DisplayPharmaceutical PreparationsPhysiological ProcessesPlayProtein BindingProtein FragmentProteinsRNARNA BindingRadialResourcesRoboticsRoleRouteScanningSeriesShapesSideSiteSite VisitSkeletonSolutionsSpecific qualifier valueStructureSubstrate SpecificitySurfaceTechniquesTimeUpdateValidationVertebral columnbasebiological researchcofactorcombinatorialcomparativecomputerized toolsdata managementdesignexperienceflexibilityfunctional groupglobular proteininstrumentmacromoleculemethod developmentmolecular dynamicsnovelnucleic acid structurepreventprogramsprotein foldingprotein structureresearch studysmall moleculetooltwo-dimensionalweb site
中文摘要
我们方法的独特性源于将蛋白质结构视为3D空间中点的集合(例如,原子坐标或描述分子表面的点),而忽略链上残基的顺序。这种基于计算机视觉和机器人的算法可以在不受顺序限制的情况下对蛋白质表面、界面或蛋白质核心进行比较。自上次实地考察以来,我们在开发新算法方面取得了实质性进展。其中一些(对接和结合位点的比较和检测)已经在上面描述过了。列举自上次现场访问以来我们开发的方法:基于残基的多蛋白质结构比较(MultiProt);蛋白质二级结构表示的多重比对(MASS);蛋白质结构在功能基团表示及其结合位点(MultiBind)和蛋白质-蛋白质界面(MAPPIS)中的多重比对;SiteEngine用于小分子和蛋白质结合位点识别,I2ISiteEngine用于界面的两两结构比较;蛋白质结构的柔性对齐(FlexProt);刚体对接(PatchDock);柔性铰链弯曲对接(FlexDock);对称对接(symdock);折叠和多分子组装组合对接(CombDock);利用噬菌体展示库(SiteLight)预测结合位点MolAxis可以高效地检测蛋白质中的通道和空洞,即使这些通道和空洞的直径非常小。此外,利用这些,蛋白质-蛋白质界面的两个非冗余数据集已经组装。这些方法都是高效的,具有最先进的功能。我已经讨论了对接方法,SiteEngine和MAPPIS(蛋白质-蛋白质接口的多重对齐)。下面我简要介绍一下FlexProt、MASS、MultiProt和MolAxis。大多数多重对准方法都是从成对对准解开始的。相比之下,MASS和MultiProt从输入分子的同时叠加中获得多个排列。此外,这两种方法都不要求所有输入分子都参与比对。实际上,它们有效地检测到输入中所有可能数量的分子的高分部分多重比对。MASS (Multiple Alignment by Secondary Structures)和MultiProt (Multiple Proteins)是全自动、高效的蛋白质结构比对检测技术,可检测输入分子之间的共同几何核心。此外,这两种方法都是序列顺序无关的。MASS基于两级对齐,同时使用二级结构和原子表示。利用二级结构信息有助于滤除噪声解,达到高效和鲁棒性。MASS能够检测非拓扑结构基序,其中二级结构以不同的顺序排列在链上。此外,MASS不仅能够检测所有输入分子共享的结构基序,还能够检测仅由分子子集共享的基序。我们已经证明了它能够处理数十个分子的顺序,检测非拓扑基序,并在输入的非预定义子集中找到具有生物学意义的排列。MASS的网址是http://bioinfo3d.cs.tau.ac.il/MASS/。MultiProt考虑用空间中的点来表示的蛋白质结构,这些点要么是c - α坐标,要么是c - α和侧链的原子或几何中心。MultiProt可在http://bioinfo3d.cs.tau.ac.il/MultiProt/上获得。我们已经说明了这两种方法在一系列应用程序中的强大功能。顺序无关性允许将MultiProt应用于结合位点和蛋白质-蛋白质界面,使MultiProt成为非常有用的结构工具。MolAxis是一个免费的,易于使用的web服务器,用于识别连接大分子外部的埋藏腔和蛋白质中的跨膜(TM)通道。生物通道对于电解质和代谢物跨膜运输和酶催化等生理过程至关重要,并且可以在底物特异性中发挥作用。由于通道识别在大分子中的重要性,我们开发了MolAxis服务器。MolAxis实现了最先进的,精确的计算几何技术,减少了通道查找问题的尺寸,使算法非常高效。给定PDB格式的蛋白质或核酸结构,服务器输出将埋藏腔连接到蛋白质外部或指向TM蛋白质中的主通道的所有可能通道。对于每个通道,门控残留物和称为“瓶颈”的最窄半径也与衬里残留物和通道表面的完整列表一起以3D图形表示。用户可以根据自己的需要操纵高级参数并指导频道搜索。MolAxis可以作为web服务器或作为独立程序在http://bioinfo3d.cs.tau.ac.il/MolAxis上获得。此外,我们一直在开发使用结构比较技术鉴定RNA非预定义三级结构的方法。我们将其应用于当前可用的RNA结构(核磁共振和晶体)的整个数据库,以获得聚类的非冗余数据集或RNA三级结构;以及识别单链RNA中挤压RNA碱基在蛋白质表面上的RNA结合位点。通道和空腔在大分子功能中发挥重要作用,是底物/产物、辅因子和药物结合、催化位点和配体/蛋白质的进出途径。此外,跨膜蛋白(TM)形成的通道作为转运体和离子通道。MolAxis是一种灵敏、快速的新工具,可用于大分子中各种尺寸和形状的通道和空腔的识别和分类。MolAxis构建了通道,这些通道代表了小分子通过通道的可能路径。分子的外中轴是具有一个以上最近原子的点的集合。它由二维表面斑块组成,可以看作是分子补体的骨架。我们在MolAxis中实现了一种新的算法,该算法使用最先进的计算几何技术来近似和扫描外内侧轴的有用子集,从而减少了问题的维度,从而使算法非常高效。MolAxis的设计目的是识别连接大分子外埋腔的通道,以及识别蛋白质中的TM通道。我们将MolAxis应用于酶腔和TM蛋白。我们进一步利用MolAxis沿着人类细胞色素P450的分子动力学轨迹监测通道尺寸。MolAxis构建了高质量的走廊,用于皮秒时间尺度间隔的快照,证实了2e衬底接入通道中的门控机制。我们将我们的结果与以前的工具在准确性、性能和寻找所需路径的基本理论保证方面进行了比较。MolAxis可以作为一个网络服务器和一个独立的易于使用的程序在网上获得(http://bioinfo3d.cs.tau.ac.il/MolAxis/)。
英文摘要
The uniqueness of our methodologies derives from viewing protein structures as collections of points (e.g., atom coordinates, or points describing molecular surfaces) in 3D space, disregarding the order of the residues on the chains. Such computer-vision and robotics-based algorithms enable comparisons of protein surfaces, interfaces, or protein cores without being limited by the sequential order. Since the last site visit, we have made substantial progress in the development of new algorithms. Some of these (docking, and binding site comparison and detection) have already been described above. To enumerate the methods we have developed since the last site visit: residue-based multiple protein structure comparison (MultiProt; multiple alignment of proteins in their secondary structure representation (MASS); multiple alignment of protein structures in the functional group representation and of their binding sites (MultiBind), and of protein-protein interfaces (MAPPIS); SiteEngine, which carries out small molecule and protein-binding site recognition and I2ISiteEngine, which carries out pairwise structural comparisons of interfaces; flexible alignment of protein structures (FlexProt; Rigid body docking (PatchDock); Flexible hinge-bending docking (FlexDock); Symmetry docking (SymmDock; Combinatorial docking for folding and multimolecular assembly (CombDock); Prediction of binding sites using phage display libraries (SiteLight); and MolAxis to detect channels and cavities in proteins in a highly efficient matter even if the diameter of these is very small. In addition, using these, two nonredundant datasets of protein-protein interfaces have been assembled. The methods are all highly efficient with state of-the-art capabilities. I have already discussed the docking methods, SiteEngine and MAPPIS (Multiple Alignment of Protein-Protein InterfaceS). Below I briefly describe FlexProt, MASS, MultiProt and MolAxis. Most methods for multiple alignment start from the pairwise alignment solutions. In contrast, MASS and MultiProt derive multiple alignments from simultaneous superpositions of input molecules. Further, both methods do not require that all input molecules participate in the alignment. Actually, they efficiently detect high scoring partial multiple alignments for all possible number of molecules in the input. MASS (Multiple Alignment by Secondary Structures) and MultiProt (Multiple Proteins) are fully automated highly efficient techniques to detect multiple structural alignments of protein structures and detect common geometrical cores between input molecules. Furthermore, both methods are sequence-order independent. MASS is based on a two-level alignment, using both secondary structure and atomic representation. Utilizing secondary structure information aids in filtering out noisy solutions and achieves efficiency and robustness. MASS is capable of detecting nontopological structural motifs, where the secondary structures are arranged in a different order on the chains. Further, MASS is able to detect not only structural motifs, shared by all input molecules, but also motifs shared only by subsets of the molecules. We have demonstrated its ability to handle on the order of tens of molecules, to detect nontopological motifs and to find biologically meaningful alignments within nonpredefined subsets of the input. MASS is available at http://bioinfo3d.cs.tau.ac.il/MASS/. MultiProt considers protein structures as represented by points in space, where the points are either the C-alpha coordinates or the C-alpha and atoms or geometric center of the side chain. MultiProt is available at http://bioinfo3d.cs.tau.ac.il/MultiProt/. We have illustrated the power of both methods on a range of applications. The order-independence allows application of MultiProt to binding sites and protein-protein interfaces, making MultiProt an extremely useful structural tool. MolAxis is a freely available, easy-to-use web server for identification of channels that connect buried cavities to the outside of macromolecules and for transmembrane (TM) channels in proteins. Biological channels are essential for physiological processes such as electrolyte and metabolite transport across membranes and enzyme catalysis, and can play a role in substrate specificity. Motivated by the importance of channel identification in macromolecules, we developed the MolAxis server. MolAxis implements state-of-the-art, accurate computational-geometry techniques that reduce the dimensions of the channel finding problem, rendering the algorithm extremely efficient. Given a protein or nucleic acid structure in the PDB format, the server outputs all possible channels that connect buried cavities to the outside of the protein or points to the main channel in TM proteins. For each channel, the gating residues and the narrowest radius termed 'bottleneck' are also given along with a full list of the lining residues and the channel surface in a 3D graphical representation. The users can manipulate advanced parameters and direct the channel search according to their needs. MolAxis is available as a web server or as a stand-alone program at http://bioinfo3d.cs.tau.ac.il/MolAxis. In addition, we have been developing methods to identify unpredefined tertiary structure of RNA using structural comparison techniques. We are applying it to the entire database of currently available RNA strucures (NMR and crystal) to derive a clustered nonredundant dataset or RNA tertiary structures; and to identify RNA binding sites on protein surfaces for extruded RNA bases in single stranded RNA. Channels and cavities play important roles in macromolecular functions, serving as access/exit routes for substrates/products, cofactor and drug binding, catalytic sites, and ligand/protein. In addition, channels formed by transmembrane (TM) proteins serve as transporters and ion channels. MolAxis is a new sensitive and fast tool for the identification and classification of channels and cavities of various sizes and shapes in macromolecules. MolAxis constructs corridors, which are pathways that represent probable routes taken by small molecules passing through channels. The outer medial axis of the molecule is the collection of points that have more than one closest atom. It is composed of two-dimensional surface patches and can be seen as a skeleton of the complement of the molecule. We have implemented in MolAxis a novel algorithm that uses state-of-the-art computational geometry techniques to approximate and scan a useful subset of the outer medial axis, thereby reducing the dimension of the problem and consequently rendering the algorithm extremely efficient. MolAxis is designed to identify channels that connect buried cavities to the outside of macromolecules and to identify TM channels in proteins. We apply MolAxis to enzyme cavities and TM proteins. We further utilize MolAxis to monitor channel dimensions along Molecular Dynamics trajectories of a human Cytochrome P450. MolAxis constructs high quality corridors for snapshots at picosecond time-scale intervals substantiating the gating mechanism in the 2e substrate access channel. We compare our results with previous tools in terms of accuracy, performance and underlying theoretical guarantees of finding the desired pathways. MolAxis is available on line as a web-server and as a stand alone easy-to-use program (http://bioinfo3d.cs.tau.ac.il/MolAxis/).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Method Development: Efficient Computer Vision Based Algo
-
批准号:7291814
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algorithms
-
批准号:8937737
-
项目类别:
-
资助金额:$10.87万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:9153571
-
项目类别:
-
资助金额:$43.97万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algorithms
-
批准号:8349006
-
项目类别:
-
资助金额:$12.85万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:8349004
-
项目类别:
-
资助金额:$64.26万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:8349005
-
项目类别:
-
资助金额:$51.4万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:8552693
-
项目类别:
-
资助金额:$53.14万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:10014370
-
项目类别:
-
资助金额:$68.71万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algorithms
-
批准号:10262089
-
项目类别:
-
资助金额:$11.83万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:10262088
-
项目类别:
-
资助金额:$47.34万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:7291812
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:8552694
-
项目类别:
-
资助金额:$42.51万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algorithms
-
批准号:8552695
-
项目类别:
-
资助金额:$10.63万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:10702352
-
项目类别:
-
资助金额:$69.45万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:7338385
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algo
-
批准号:7338445
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Method Development: Efficient Computer Vision Based Algorithms
-
批准号:8763103
-
项目类别:
-
资助金额:$9.89万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Biomolecular Recognition and Binding Mechanisms
-
批准号:7733032
-
项目类别:
-
资助金额:$67.32万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:7592701
-
项目类别:
-
资助金额:$57.07万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
Protein Structure, Stability, and Amyloid Formation
-
批准号:10262087
-
项目类别:
-
资助金额:$59.17万
-
财政年份:--
-
负责人:Ruth Nussinov
-
依托单位:
海外基金