FRAGS: estimation of coding sequence substitution rates from fragmentary data.

FRAGS: estimation of coding sequence substitution rates from fragmentary data.
复制标题

DOI:
10.1186/1471-2105-5-8
复制
发表时间:
2004-01-29
期刊:
影响因子:
3
通讯作者:
Seoighe C
Seoighe C
中科院分区:
生物学4区
文献类型:
--
作者:
Swart EC;Hide WA;Seoighe C

文献摘要

参考文献

被引文献

相似文献

蛋白质编码序列中的替换率可以为生物医学和理论上感兴趣的进化过程提供重要的见解。编码序列数据的增加使研究人员能够更准确地估计生物体对的编码序列分歧。然而,使用不同的数据源、比对方案和方法来估计替换率导致对定义直系基因编码序列分歧的关键参数的估计存在很大差异。尽管并非所有生物都有完整的基因组序列数据,但片断序列数据可以提供对替代率的准确估计,只要使用适当和一致的方法,并考虑到从不同数据来源获得的估计的差异。我们已经开发了FRAGS,这是一个应用程序框架,它使用现有的、可免费获得的软件组件来构建帧内比对,并从零碎的序列数据中估计编码替换率。对由FRAGS产生的人类和黑猩猩序列的编码序列替换估计表明,方法上的差异可能导致对重要替换参数的显著不同估计。估计的替换率也被用来推断我们所分析的数据集中测序误差的上限。我们已经开发了一种系统,可以对来自一对有机体的同源序列的替换率进行稳健的估计。我们的系统可以用于从其中一个生物体获得零碎的基因组或转录本数据,而另一个是Ensambl数据库中的完全测序的基因组。除了估计替换统计数据外,我们的系统还使用户能够管理和查询比对和替换数据。
Rates of substitution in protein-coding sequences can provide important insights into evolutionary processes that are of biomedical and theoretical interest. Increased availability of coding sequence data has enabled researchers to estimate more accurately the coding sequence divergence of pairs of organisms. However the use of different data sources, alignment protocols and methods to estimate substitution rates leads to widely varying estimates of key parameters that define the coding sequence divergence of orthologous genes. Although complete genome sequence data are not available for all organisms, fragmentary sequence data can provide accurate estimates of substitution rates provided that an appropriate and consistent methodology is used and that differences in the estimates obtainable from different data sources are taken into account. We have developed FRAGS, an application framework that uses existing, freely available software components to construct in-frame alignments and estimate coding substitution rates from fragmentary sequence data. Coding sequence substitution estimates for human and chimpanzee sequences, generated by FRAGS, reveal that methodological differences can give rise to significantly different estimates of important substitution parameters. The estimated substitution rates were also used to infer upper-bounds on the amount of sequencing error in the datasets that we have analysed. We have developed a system that performs robust estimation of substitution rates for orthologous sequences from a pair of organisms. Our system can be used when fragmentary genomic or transcript data is available from one of the organisms and the other is a completely sequenced genome within the Ensembl database. As well as estimating substitution statistics our system enables the user to manage and query alignment and substitution data.
DOI: 10.1073/pnas.96.8.4482
发表时间: 1999-04-13
影响因子: 11.1
作者:
Duret, L;Mouchiroud, D
通讯作者: Mouchiroud, D
DOI: 10.1093/nar/22.12.2360
发表时间: 1994-06-25
影响因子: 14.9
作者:
DURET, L;MOUCHIROUD, D;GOUY, M
通讯作者: GOUY, M
DOI: 10.1101/gr.212002
发表时间: 2002-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lee, Y;Sultana, R;Quackenbush, J
通讯作者: Quackenbush, J
DOI: 10.1093/nar/30.1.38
发表时间: 2002-01-01
影响因子: 14.9
作者:
Hubbard, T;Barker, D;Clamp, M
通讯作者: Clamp, M
DOI: 10.1126/science.1080600
发表时间: 2003-04-11
期刊: SCIENCE
影响因子: 56.9
作者:
Navarro, A;Barton, NH
通讯作者: Barton, NH