课题基金 / 基金详情

Statistics of Sequence Comparison

Statistics of Sequence Comparison
序列比较统计
批准号:
8149590
负责人:
STEPHEN F ALTSCHUL
金额:
$39.17万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
关键词:

项目摘要

项目成果

STEPHEN F ALTSCHUL的其他基金

相似基金

相关文献

中文摘要
翻译
今年的工作重点是多个项目的评分系统 对齐。最多配对和多序列比对 程序寻求具有最佳分数的对齐。中心到 定义这样的分数就是选择一组替补 对齐的氨基酸或核苷酸的分数。替代 局部配对的得分隐含为 对数赔率表,比较对齐的概率 关联性和非关联性模式下的两个字母, 并且最佳成对替换分数是明确的 就是这样建造的。我们已经开发了一些想法,基于 最小描述长度原则,用于扩展此 形式主义到多重结盟。最简单的,贝叶斯 可以使用各种方法来得出“Bild”替补分数 来自描述相关列的先前分发 信件。这种方法以前只被使用过 定义将单个序列比对到 序列图谱,但它有更广泛的适用性。 我们在Gibbs抽样中使用了《图片报》的分数 优化过程,并表明它们产生了 改进了生物构建方面的性能 精确对齐。我们已经制定了试点计划 用于使用Bild构建有间隙的多条路线 得分。使用人工序列,我们已经展示了这些 比早期程序具有更高性能的程序 在检测结构域边界方面。我们也为他们鼓掌 对DNA结合域的识别和注释 在Apicomplexan蛋白中。 在相关工作中,我们研究了Dirichlet混合模型 在非标准序列组成的背景下。一个 字母表上具有M分量的Dirichlet混合词 L字母有M*(L 1)-1自由参数。如果M=L/2,则这 与对称成对替换的数量完全相同 矩阵。虽然每个狄利克雷特混合物都暗示着一个独特的 这样的矩阵,我们已经证明了多个混合物可以映射 相同的矩阵,而某些替换矩阵可能不同 对应于任何狄里克莱特混合物。狄里克莱特混合物 用于蛋白质序列分析的一般结构 来自一组特定的蛋白质,意味着特定的 背景氨基酸组成。混合物应该是 对于蛋白质与蛋白质的比较不是最佳的 明显不同的成分。我们已经描述了 一种明智而有效的方法来调整 Dirichlet混合物的参数,所以它们是 与任何特定成分一致。
英文摘要
Work this year focused on scoring systems for multiple alignment. Most pairwise and multiple sequence alignment programs seek alignments with optimal scores. Central to defining such scores is selecting a set of substitution scores for aligned amino acids or nucleotides. Substitution scores for local pairwise alignment are implicitly of log-odds form, comparing the probabilities of aligning two letters under models of relatedness and non-relatedness, and the best pairwise substitution scores are explicitly so constructed. We have developed ideas, based on the minimum description length principle, for extending this formalism to multiple alignments. Most simply, Bayesian methods can be used to derive "BILD" substitution scores from prior distributions describing columns of related letters. This approach has been used previously only to define scores for aligning individual sequences to sequence profiles, but it has much broader applicability. We have employed BILD scores in Gibbs sampling optimization procedures, and shown that they yield improved performance in constructing biologically accurate alignments. We have developed pilot programs for constructing gapped multiple alignments using BILD scores. Using artificial sequences, we have shown these programs to have superior performance to earlier programs at detecting domain boundaries. We have also appled them to the recognition and annotation of DNA-binding domains in Apicomplexan proteins. In related work, we have studied Dirichlet mixture models in the context of non-standard sequence composition. A Dirichlet mixture with M components over an alphabet of L letters has M*(L+1)-1 free parameters. If M = L/2, this is exactly as many as a symmetric pairwise substitution matrix. While each Dirichlet mixture implies a unique such matrix, we have shown that multiple mixtures can map to the same matrix, and some substitution matrices may not correspond to any Dirichlet mixture. A Dirichlet mixture for protein sequence analysis generally is constructed from a particular set of proteins, implying a particular background amino acid composition. The mixture should be non-optimal for the comparison of proteins with significantly different composition. We have described a sensible and efficient method for adjusting the parameters of a Dirichlet mixture so that they are consistent with any specified composition.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
STATISTICS OF SEQUENCE COMPARISON
  • 批准号:
    6290478
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN F ALTSCHUL
  • 依托单位:
Improvements And Extensions To The Blast Algorithms
  • 批准号:
    6546809
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN F ALTSCHUL
  • 依托单位:
Improvements And Extensions To The Blast Algorithms
  • 批准号:
    6843572
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    STEPHEN F ALTSCHUL
  • 依托单位:
Statistics of Sequence Comparison
  • 批准号:
    9160904
  • 项目类别:
  • 资助金额:
    $20.45万
  • 财政年份:
    --
  • 负责人:
    STEPHEN F ALTSCHUL
  • 依托单位:
海外基金