A comprehensive system for consistent numbering of HCV sequences, proteins and epitopes

A comprehensive system for consistent numbering of HCV sequences, proteins and epitopes
复制标题

DOI:
10.1002/hep.21377
复制
发表时间:
2006-11-01
期刊:
影响因子:
13.5
通讯作者:
Simmonds, Peter
Simmonds, Peter
中科院分区:
医学1区
文献类型:
--
作者:
Kuiken, Carla;Combet, Christophe;Simmonds, Peter

文献摘要

被引文献

相似文献

相关病毒讨论HCV序列数据库如何引入和促进HCV核苷酸、蛋白质和表位的标准化编号系统。DNA和蛋白质序列中不一致和不准确的位置编号是HCV科学文献中的一个问题。在HCV的功能和临床研究中,越来越需要编号的一致性。例如,一种明确的方法来参考NS3和NS5B编码序列中与特定HCV抑制剂耐药性相关的特定位置的氨基酸取代,在抗病毒治疗的研究中是必不可少的。本文提供了一个实用指南,以帮助在将来规避这些问题,并在该领域的讨论中引入一种通用语言。目前系统的范围仅限于HCV多蛋白和非翻译区(UTRs);由于替代阅读框蛋白的争议性和极端长度变化,如果需要,这些蛋白的编号将在以后决定。我们提出了一个编号系统改编自洛斯阿拉莫斯艾滋病毒数据库,1与元素从乙型肝炎病毒编号系统。该系统包括核苷酸和氨基酸序列以及表位。它使用分离物H77的全长基因组序列(登录号AF009606)作为参考,并包括相对于该参考序列的插入和缺失编号方法。选择H77是因为它是许多不同种类功能研究的常用参考菌株。此外,该序列的RNA转录本具有传染性,证明该序列的5端和3端是完整的。表1列出了HCV基因组区域的边界,图1提供了完整的AF009606 HCV基因组序列的详细核酸和氨基酸编号。
Related Viruses to discuss how HCV sequence databases could introduce and facilitate a standardized numbering system for HCV nucleotides, proteins and epitopes. Inconsistent and inaccurate numbering of locations in DNA and protein sequences is a problem in the HCV scientific literature. Consistency in numbering is increasingly required for functional and clinical studies of HCV. For example, an unambiguous method for referring to amino acid substitutions at specific positions in NS3 and NS5B coding sequences associated with resistance to specific HCV inhibitors is essential in the investigation of antiviral treatment. This article provides a practical guide to help circumvent these problems in the future, and to bring a common language into discussions in the field. The scope of the current system is limited to the HCV polyprotein and the untranslated regions (UTRs); because of the controversial nature and extreme length variation of the alternate reading frame proteins, numbering for these proteins, if needed, will be decided at a later date.We propose a numbering system adapted from the Los Alamos HIV database, 1 with elements from the hepatitis B virus numbering system. 2 The system comprises both nucleotides and amino acid sequences and epitopes. It uses the full length genome sequence of isolate H77 (accession number AF009606) as a reference, and includes a method for numbering insertions and deletions relative to this reference sequence. H77 was chosen because it is a commonly used reference strain for many different kinds of functional studies. Furthermore, RNA transcripts from this sequence are of demonstrated infectivity, 3, 4 providing evidence that the 5 and 3 ends of the sequence are complete. Table 1 lists the boundaries of HCV genomic regions and Fig. 1 provides detailed nucleic acid and amino acid numbering over the complete AF009606 HCV genome sequence.