A comprehensive system for consistent numbering of HCV sequences, proteins and epitopes
A comprehensive system for consistent numbering of HCV sequences, proteins and epitopes
复制标题
DOI:
10.1002/hep.21377
复制
发表时间:
2006-11-01
期刊:
影响因子:
13.5
通讯作者:
Simmonds, Peter
中科院分区:
文献类型:
--
作者:
Kuiken, Carla;Combet, Christophe;Simmonds, Peter
Related Viruses to discuss how HCV sequence databases could introduce and facilitate a standardized numbering system for HCV nucleotides, proteins and epitopes. Inconsistent and inaccurate numbering of locations in DNA and protein sequences is a problem in the HCV scientific literature. Consistency in numbering is increasingly required for functional and clinical studies of HCV. For example, an unambiguous method for referring to amino acid substitutions at specific positions in NS3 and NS5B coding sequences associated with resistance to specific HCV inhibitors is essential in the investigation of antiviral treatment. This article provides a practical guide to help circumvent these problems in the future, and to bring a common language into discussions in the field. The scope of the current system is limited to the HCV polyprotein and the untranslated regions (UTRs); because of the controversial nature and extreme length variation of the alternate reading frame proteins, numbering for these proteins, if needed, will be decided at a later date.We propose a numbering system adapted from the Los Alamos HIV database, 1 with elements from the hepatitis B virus numbering system. 2 The system comprises both nucleotides and amino acid sequences and epitopes. It uses the full length genome sequence of isolate H77 (accession number AF009606) as a reference, and includes a method for numbering insertions and deletions relative to this reference sequence. H77 was chosen because it is a commonly used reference strain for many different kinds of functional studies. Furthermore, RNA transcripts from this sequence are of demonstrated infectivity, 3, 4 providing evidence that the 5 and 3 ends of the sequence are complete. Table 1 lists the boundaries of HCV genomic regions and Fig. 1 provides detailed nucleic acid and amino acid numbering over the complete AF009606 HCV genome sequence.