The MEME suite of motif-based sequence analysis tools
The MEME suite of motif-based sequence analysis tools
批准号:
9040997
负责人:
Timothy L Bailey
金额:
$35.66万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-28 至 2018-03-31
关键词:
Active SitesAlgorithmsAmino Acid MotifsAmino Acid SequenceBase SequenceBinding ProteinsBinding SitesBiologicalBiological ModelsCell physiologyCellsCharacteristicsCollectionComplexComputer softwareDNADNA SequenceDataData SetDatabasesDiseaseDisease modelEnsureEnzymesEvolutionGene ExpressionGenesGenomeGenomic DNAGenomicsGoalsHealthIndividualKnowledgeLinuxMass Spectrum AnalysisMediatingMicroRNAsModelingMolecularMolecular BiologyMolecular EvolutionNatureOutputPaperPatternPeptide Sequence DeterminationPeptide Signal SequencesPhosphorylationPhosphotransferasesPositioning AttributePost Translational Modification AnalysisPost-Translational Protein ProcessingProceduresProcessProteinsProteomicsRNARNA SequencesRNA SplicingResearch InfrastructureRunningScanningScientistSequence AnalysisSeriesSoftware EngineeringStatistical ModelsSystemTestingTimeWorkcell typechromatin immunoprecipitationcomputerized toolsdesignexperiencehuman diseaseimprovedinsightinteroperabilitymodel buildingmolecular scaleprotein foldingprotein structuresoftware developmenttoolusabilityweb portal
中文摘要
描述(申请人提供):这个项目的广泛目标是开发和应用计算机工具来检测、建模和理解生物上重要的序列模式,称为基序,在基因组、RNA和蛋白质中编码。序列基序携带了许多细胞正常运作所必需的信息。例如,基因组DNA中的基序包含有助于调节基因表达的信息。RNA中的序列基序编码剪接和调控信息,如microRNA结合位点。在蛋白质水平上,序列基序可以参与酶结合位点,为蛋白质结构提供锚定,或介导翻译后修饰,如由激酶进行的磷酸化。我们使用统计模型对生物序列模式进行建模,该模型捕捉局部序列模式,同时允许自然发生的可变性。自2011年以来,已有33,000多名独立用户访问了Meme Suite门户网站,而且用户数量一直在稳步增长。根据谷歌学者的数据,截至2013年6月28日,描述表情包套件的论文已被引用6827次。在提出的项目中,我们的目标是改进Meme Suite中的核心算法,为Suite添加重要的新功能,并提高软件的健壮性、可靠性和可用性。特别是,我们将显著提高
核心基序发现算法能够扩展到更大的数据集,识别新类型的基序,并提供更准确的统计置信度估计。我们将为该套件添加功能,以允许用户识别和表征与翻译后蛋白质修饰相关的基序。我们还将进行一系列软件工程和可用性改进,这些改进将极大地提高整体用户体验。我们的软件可以在本地安装,也可以通过我们的门户网站远程运行,以对大型、复杂的基因组和蛋白质组数据集执行各种不同的分析。它被世界各地的科学家广泛使用。我们的目标是继续维护和开发这个软件,促进科学发现,并导致对分子生物学和人类疾病的广泛基本过程的洞察。
英文摘要
DESCRIPTION (provided by applicant): The broad goal of this project is to develop and apply computational tools for detecting, modeling and understanding biologically important sequence patterns, called motifs, encoded in the genome, in RNA and in proteins. Sequence motifs carry much of the information essential to the correct functioning of cells. For ex- ample, motifs in genomic DNA contain information that helps to regulate gene expression. Sequence motifs in RNA encode splice junctions and regulatory information such as microRNA binding sites. At the protein level, sequence motifs may participate in enzymatic binding sites, provide anchors for a protein structure or mediate posttranslational modifications such as phosphorylation by kinases. We model biological sequence patterns using statistical models that capture local sequence patterns while allowing for naturally occurring variability. Since 2011 more than 33,000 unique users have accessed the MEME Suite web portal, and the number of users has been steadily growing. As of June 28, 2013, the papers describing the MEME Suite have been cited 6827 times, according to Google scholar. In the proposed project, we aim to improve the core algorithms in the MEME Suite, add significant new functionality to the Suite, and improve the robustness, reliability and usability of software. In particular, we will significantly enhance the
core motif discovery algorithm to scale to larger data sets, to identify new types of motifs, and t provide more accurate statistical confidence estimates. We will add functionality to the Suite to allow users to identify and characterize motifs associated with post-translational protein modifications. We will also carry out a series of software engineering and usability improvements that will greatly enhance the overall user experience. Our software can be locally installed or run remotely through our web portal to perform a diverse set of analyses on large, complex genomic and proteomic data sets. It is in widespread use by scientists around the world. We aim to continue to maintain and develop this software, facilitating scientific discovery and leading to insights into a wide spectrum of fundamental processes in molecular biology and human disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
The MEME suite of motif-based sequence analysis tools
-
批准号:9251869
-
项目类别:
-
资助金额:$35.66万
-
财政年份:2009
-
负责人:Timothy L Bailey
-
依托单位:
海外基金