课题基金 / 基金详情

A CRITICAL EVALUATION AND COMPARISON OF COMPUTERIZED SEQUENCE ANALYSIS PROGRAMS

A CRITICAL EVALUATION AND COMPARISON OF COMPUTERIZED SEQUENCE ANALYSIS PROGRAMS
计算机化序列分析程序的批判性评估和比较
批准号:
3752825
负责人:
J I POWELL
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

J I POWELL的其他基金

相关文献

中文摘要
翻译
与M博士合作。米勒,NCI,一个关键的,定量的 分析了几种商业序列的组装和分析 包装件.当代分子生物学的一个基本问题是 DNA序列的测定和解释。由于局限性, 目前的测序技术,序列确定需要拼接 短的、重叠的序列片段结合在一起, 连续序列许多商业计算机程序已经被 自动化这个过程。虽然对个别一揽子计划的审查 已经发表,这是第一个已知的研究, 比较了这些程序的装配精度。 11个项目被选中,主要是基于他们的 在NIH校园内的可用性。序列数据不是随机的,而是包含 有序的重复序列。同样,测序测定中的错误 并不是随机分布的。为了提供一个可控的, 用于测量性能和准确性的真实数据集, 序列,大鼠多药耐药基因(RATMDRM,5254个碱基对, 登录号M62425)分成58个随机重叠片段, 200到400个碱基对。然后将这些随机接种0至 基于最初使用的片段的错误分布的15%错误 以确定序列。错误的形式是错误的基地, 删除碱基或添加碱基。 测试的程序根据准确性分为三大类。在 为了排除所选测试序列特有的条件, 4500 - 4600碱基对的其它序列用于重复 测试除了一个例外, 使用RATMDRM时遇到。此外,还测试了一些程序, RATMDRM的不同排列,以确定他们的能力, 组装序列而不管片段的输入顺序。 还比较了编辑组装序列的容易性。成果 这项研究被《生物学杂志》接受发表。 计算。
英文摘要
In collaboration with Dr. M. Miller, NCI, a critical, quantitative analysis was done of several commercial sequence assembly and analysis packages. A fundamental problem in contemporary molecular biology is the determination and interpretation of DNA sequences. Due to limitations of current sequencing technology, sequence determination entails the piecing together of short, overlapping sequence fragments into a single, long contiguous sequence. A number of commercial computer programs have been marketed to automate this process. While reviews of individual packages have been published, this is the first known study that critically compares the accuracy of assembly by these programs. Eleven programs were selected, primarily on the basis of their availability on the NIH campus. Sequence data is not random, but contains ordered repeated sequences. Likewise, errors in sequencing determinations are not randomly distributed. In order to provide a controlled and realistic dataset for measuring performance and accuracy, a known sequence, the rat multidrug resistance gene (RATMDRM, 5254 base pairs, accession number M62425) was split into 58 random overlapping fragments of 200 to 400 base pairs in length. These were then randomly seeded with 0 to 15% error based on the error distribution of the fragments originally used to determine the sequence. Errors were in the form of miscalled bases, deleted bases or added bases. The programs tested fell into three general groups based on accuracy. In order to rule out conditions unique to the chosen test sequence, four other sequences of between 4500 and 4600 base pairs were used to repeat the tests. With one exception, the error rates were comparable to those encountered using RATMDRM. Additionally, some programs were tested with different permutations of RATMDRM to ascertain their capacity to properly assemble the sequence regardless of the order of input of the fragments. Ease of editing the assembled sequences was also compared. Results of this study were accepted for publication by the Journal of Biological Computation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
UTILIZATION OF SPECIALIZED HARDWARE FOR DNA SEQUENCE ANALYSIS
LABORATORY ANALYSIS PACKAGE
UTILIZATION OF SPECIALIZED HARDWARE FOR DNA SEQUENCE ANALYSIS
UTILIZATION OF SPECIALIZED HARDWARE FOR DNA SEQUENCE ANALYSIS