Software development and application of a simulation framework for protein evolution
Software development and application of a simulation framework for protein evolution
批准号:
8835870
负责人:
Stephanie J Spielman
金额:
$4.12万
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-06-01 至 2016-05-31
关键词:
AccountingAddressAdoptedAdoptionAminesAmino Acid SequenceAmino AcidsAntiviral AgentsArchitectureBehaviorBiologyBiomedical ResearchCodeCodon NucleotidesCommunitiesComputational BiologyComputer SimulationComputer softwareDataDevelopmentDiseaseDisease OutbreaksEnsureEventEvolutionFoundationsFutureGeneticGoalsHereditary DiseaseHeterogeneityLiteratureMethodologyMethodsModelingMutationNatural SelectionsPeptide Sequence DeterminationPhylogenetic AnalysisPhylogenyPositioning AttributeProcessProtein DynamicsProteinsResearchResearch PersonnelRoleSequence AnalysisShapesSiteSoftware DesignSolidTertiary Protein StructureTestingTimeVaccine DesignVirulenceanalytical toolbasebiological researchcareerclinically relevantcomparativecomputing resourcesflexibilityinsertion/deletion mutationinsightopen sourcepreferencepublic health relevancesimulationsimulation softwaresoftware developmenttooltool developmentuser friendly softwareuser-friendly
中文摘要
描述(由申请人提供):描述蛋白质编码序列进化动力学的方法是比较序列分析中最广泛使用的工具之一,其应用范围从识别关键功能蛋白残基到预测疾病的进化轨迹和毒性。传统的蛋白质编码序列进化模型侧重于确定蛋白质的进化速率,或者蛋白质中不同位置的进化速度。然而,尽管这些模型被广泛应用,但它们忽略了蛋白质进化动力学的一个关键方面:自然选择倾向于氨基酸在蛋白质中不同位置的独特、位点特异性分布。传统模型忽略了这一总体约束,只评估蛋白质氨基酸是否发生了变化。为了解决这一限制,出现了一类被称为“突变选择”的模型,它明确地解释了氨基酸偏好的影响。尽管突变选择模型是在15年前首次提出的,但其高昂的计算成本限制了其应用。然而,在过去的一年里,计算能力的提高使这些模型第一次变得易于处理。随着计算能力的进一步提高,突变选择模型将不断发展,并在序列分析研究中发挥核心作用。因此,科学界将需要一套工具来评估和检验关于这些模型的假设的有效性。为此,我将开发软件,根据突变选择模型沿系统发育模拟蛋白质编码序列。遗传数据的模拟是一种广泛使用的方法来验证和比较分析工具,但没有可用的序列模拟软件考虑突变选择模型。为此,我将开发一种高度灵活、用户友好的工具,并将其传播给科学界
英文摘要
DESCRIPTION (provided by applicant): Methods characterizing the evolutionary dynamics of protein-coding sequences are among the most widely-used tools in comparative sequence analysis, with applications ranging from identifying key functional protein residues to predicting the evolutionary trajectories and virulence of disease. Traditional models of protein-coding sequence evolution focus on identifying protein evolutionary rates, or how quickly different positions in a protein evolve. However, while widely implemented, such models overlook a key aspect of protein evolutionary dynamics: natural selection favors distinct, site-specific distributions of amino acids across positions in proteins. Traditional models ignore this overarching constraint and assess only whether protein amino acids change. To address this limitation, a class of models known as "mutation-selection" models, which explicitly account for the effects of amino acid preferences, have emerged. Although mutation-selection models were first proposed over 15 years ago, their high computational expense has limited their use. However, within the past year, increases in computational power have made these models tractable for the first time. As this computational power increases further, it is clear that mutation-selection models will progress and take a central role in sequence analysis studies. The scientific community will therefore need a set of tools which can assess the validity of and test hypotheses regarding these models. To this end, I will develop software to simulate protein-coding sequences along phylogenies according to mutation-selection models. Simulation of genetic data is a widely-used approach to verify and compare analytical tools, but there is no available sequence-simulation software which considers mutation-selection models. I will develop a highly flexible, user-friendly tool for this purpose and disseminate it to the scientific
community. My software will incorporate realistic protein dynamics, including heterogeneity, domains, and insertion and deletion events, into simulations. Subsequently, I will use this tool to
conduct a comprehensive comparison between the two available mutation-selection model inference methods. These recently introduced methods produce distinct, incompatible results, and as a consequence it remains unclear which method is preferred for sequence analysis. I will systematically examine the limitations and capabilities of each model to reveal under which conditions each model is preferred. This study will provide valuable guidance to researchers in selecting robust methodologies.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
海外基金