课题基金 / 基金详情

ABI Innovation: New methods for multiple sequence alignment with improved accuracy and scalability

ABI Innovation: New methods for multiple sequence alignment with improved accuracy and scalability
ABI Innovation:多序列比对的新方法,具有更高的准确性和可扩展性
批准号:
1458652
负责人:
Tandy Warnow
金额:
$86.16万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-15 至 2021-07-31

项目摘要

项目成果

Tandy Warnow的其他基金

相似基金

相关文献

中文摘要
翻译
多序列比对(MSA)是最基本的生物信息学步骤之一,其中一组分子序列(即,DNA、RNA或氨基酸序列)排列在矩阵内以识别相应的位置。MSA计算是许多生物分析中基本的第一步。由于其广泛的适用性和重要性,许多MSA方法已被开发并广泛使用。 不幸的是,许多真实的世界生物数据集具有使得精确的MSA计算非常困难的特征(例如,大尺寸和片段序列)。 由于估计不佳的比对导致下游生物分析中的错误,因此需要新的MSA技术,该技术可以在困难的数据集上产生准确的比对。该项目将开发MSA方法,大大提高准确性,可以分析在全国不同生物学项目中组装的大型和异构序列数据集。该项目也有一个实质性的外展组成部分,女子学院和少数民族服务机构,暑期软件学校,以培训生物学家在使用项目软件。多序列比对(MSA)和同源性估计是两个非常基本的生物信息学问题,这坐在机器学习,统计估计,进化和结构生物学的交叉点。MSA在构建进化树,理解蛋白质的功能和结构,检测蛋白质之间的相互作用,甚至基因组组装方面具有特别重要的意义。大规模MSA和概率估计也需要高性能计算和并行算法,以便提供足够的可扩展性。该团队将开发新的机器学习技术,以极大地改进MSA方法,从而改进遗传学估计,因为它依赖于准确的多序列比对。该项目的核心是算法开发,利用各种机器学习技术(包括隐马尔可夫模型),统计估计方法(特别是贝叶斯MCMC和最大似然)和新颖的算法策略,所有这些都专注于提高可扩展性和准确性。有关该项目的更多信息,请访问:http://tandy.cs.illinois.edu/MSAproject.html
英文摘要
Multiple sequence alignment (MSA) is one of the most basic bioinformatics steps, in which a set of molecular sequences (i.e., DNA, RNA, or amino acid sequences) are arranged inside a matrix to identify corresponding positions. MSA calculation is a fundamental first step in many biological analyses. Because of its broad applicability and importance, many MSA methods have been developed and are in wide use today. Unfortunately, many real world biological datasets have features (large size and fragmentary sequences, for example) that make accurate MSA calculation very difficult. Because poorly estimated alignments result in errors in downstream biological analyses, new MSA techniques are needed that can produce accurate alignments on difficult datasets. This project will develop MSA methods with greatly improved accuracy, and that can analyze the large and heterogeneous sequence datasets being assembled in different biology projects nationally. The project also has a substantial outreach component to women's colleges and minority serving institutions, and summer software schools to train biologists in the use of the project software.Multiple sequence alignment (MSA) and phylogeny estimation are two very basic bioinformatics problems, which sit at the intersection of machine learning, statistical estimation, and evolutionary and structural biology. MSA has particular importance in constructing evolutionary trees, understanding the function and structure of proteins, detecting interactions between proteins, and even genome assembly. Large-scale MSA and phylogeny estimation also require high performance computing and parallel algorithms, in order to provide adequate scalability. The team will develop new machine learning techniques to greatly improve MSA methods, and hence also phylogeny estimation, since it depends on accurate multiple sequence alignments. The core of this project is algorithm development, utilizing a variety of machine learning techniques (including Hidden Markov Models), statistical estimation methods (especially Bayesian MCMC and maximum likelihood), and novel algorithmic strategies, all focused on improving scalability and accuracy. More information about the project can be found at: http://tandy.cs.illinois.edu/MSAproject.html
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1093/bioinformatics/btab788
发表时间: 2022-01-27
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Shen C, Zaharias P, Warnow T]
通讯作者: Warnow T
IIBR Informatics: Advancing Bioinformatics Methods using Ensembles of Profile Hidden Markov Models
AitF: Full: Collaborative Research: Graph-theoretic algorithms to improve phylogenomic analyses
III: AF: Medium: Collaborative Research: Scalable and Highly Accurate Methods for Metagenomics
Collaborative Research: Novel Methodologies for Genome-scale Evolutionary Analysis of Multi-locus data
海外基金