Evading the annotation bottleneck: using sequence similarity to search non-sequence gene data

Evading the annotation bottleneck: using sequence similarity to search non-sequence gene data
复制标题

DOI:
10.1186/1471-2105-9-442
复制
发表时间:
2008-10-17
期刊:
影响因子:
3
通讯作者:
Papalopulu, Nancy
Papalopulu, Nancy
中科院分区:
生物学4区
文献类型:
--
作者:
Gilchrist, Michael J.;Christensen, Mikkel B.;Papalopulu, Nancy

文献摘要

被引文献

相似文献

背景:非序列基因数据(图像、文献等)可以在许多不同的公共数据库中找到。对这些数据的访问主要是通过使用基因名称的基于文本的方法;然而,生物之间的基因注释既不完整,也不完全系统,而且随着时间的推移也不稳定。这给基于文本的访问带来了一些挑战,尤其是跨物种搜索。提出了一种基于序列相似性的非序列数据检索方法,消除了对标注和文本搜索的依赖。这项工作的动机是需要提供更好地访问大量的原位图像,并观察到这些图像数据通常与特定的基因序列有关。序列相似性搜索可以在现有的面向基因的数据库中找到,但大多数是通过导航链接间接访问非序列数据。结果:构建了三个应用程序来探索所提出的方法:访问图像数据、文献和基因名称。搜索从用户感兴趣的基因序列开始,然后根据与目标数据相关的序列数据库进行搜索。匹配的(非序列的)目标数据直接返回到用户的浏览器,按序列相似性组织。该方法很好地满足了图像数据管理的预期应用。与基于文本的图像数据集搜索结果的比较表明了该方法的准确性。应用于文献检索,它促进了大多数高相关性参考文献的检索。将其应用于基因名称数据,为物种内和物种间相关基因的名称变异提供了有益的分析。结论:该方法为现有的基于文本检索或编辑基因列表的基因数据检索方法提供了一个强大而有用的补充。特别地,该方法促进了跨物种比较,并且能够处理新的或其他未注释的基因。使用该方法的应用程序快速且易于构建,并且数据几乎不需要维护。这种方法在很大程度上避免了对注释的需要,而注释可能是开发基因组规模数据资源的主要障碍。
Background: Non-sequence gene data (images, literature, etc.) can be found in many different public databases. Access to these data is mostly by text based methods using gene names; however, gene annotation is neither complete, nor fully systematic between organisms, and is also not generally stable over time. This provides some challenges for text based access, especially for cross-species searches. We propose a method for non-sequence data retrieval based on sequence similarity, which removes dependence on annotation and text searches. This work was motivated by the need to provide better access to large numbers of in situ images, and the observation that such image data were usually associated with a specific gene sequence. Sequence similarity searches are found in existing gene oriented databases, but mostly give indirect access to non-sequence data via navigational links.Results: Three applications were built to explore the proposed method: accessing image data, literature and gene names. Searches are initiated with the sequence of the user's gene of interest, which is searched against a database of sequences associated with the target data. The matching (non-sequence) target data are returned directly to the user's browser, organised by sequence similarity. The method worked well for the intended application in image data management. Comparison with text based searches of the image data set showed the accuracy of the method. Applied to literature searches it facilitated retrieval of mostly high relevance references. Applied to gene name data it provided a useful analysis of name variation of related genes within and between species.Conclusion: This method makes a powerful and useful addition to existing methods for searching gene data based on text retrieval or curated gene lists. In particular the method facilitates cross-species comparisons, and enables the handling of novel or otherwise un-annotated genes. Applications using the method are quick and easy to build, and the data require little maintenance. This approach largely circumvents the need for annotation, which can be a major obstacle to the development of genomic scale data resources.