Establishing a method of vector contamination identification in database sequences

Establishing a method of vector contamination identification in database sequences
复制标题

DOI:
10.1093/bioinformatics/15.2.106
复制
发表时间:
1999-02-01
期刊:
影响因子:
5.8
通讯作者:
Schad, PA
Schad, PA
中科院分区:
生物学3区
文献类型:
--
作者:
Seluja, GA;Farmer, A;Schad, PA

文献摘要

被引文献

相似文献

动机:核苷酸序列数据库是私人和学术研究团体的宝贵工具,从序列检索到同源性搜索。数据库面临着与数据质量相关的几个问题,例如测序伪像和错误的存在。我们调查了这些错误的一个主要来源,即存在的载体污染sequence.Results:使用面板的180载体polylinker序列,我们发现0.36%或3029载体匹配序列在GenBank Release 95-96中,与平均载体匹配:长度为72个核苷酸。载体污染序列的数量随着数据库的增长而增长;然而,从1982年到1996年,污染百分比保持在平均0.28%左右。
Motivation: The nucleotide sequence databases are invaluable tools both for the private and the academic research communities, from the retrieval of sequences to homology searching. Several issues related to data quality, such as the existence of sequencing artifacts and errors, are facing the databases. We investigated a major source of these errors, i.e. the presence of vector-contaminated sequences.Results: Using a panel of 180 vector polylinker sequences, we found 0.36% or 3029 vector-matching sequences in GenBank Release 95-96, with an average vector-matching: length of 72 nucleotides. The number of vector-contaminated sequences has been growing with the database; however, the percent contamination has remained approximately constant at an average of 0.28% from 1982 to 1996.