Evaluating annotations of an Agilent expression chip suggests that many features cannot be interpreted

Evaluating annotations of an Agilent expression chip suggests that many features cannot be interpreted
复制标题

DOI:
10.1186/1471-2164-10-566
复制
发表时间:
2009-11-30
期刊:
影响因子:
4.4
通讯作者:
Schaeffer, Alejandro A.
Schaeffer, Alejandro A.
中科院分区:
生物学2区
文献类型:
--
作者:
Gertz, E. Michael;Sengupta, Kundan;Schaeffer, Alejandro A.

文献摘要

被引文献

相似文献

背景资料:在尝试重新分析来自Agilent 4 x 44人类表达芯片的已发表数据时,我们发现一些60聚体寡核苷酸特征不能被解释为代表单个人类基因。例如,一些寡核苷酸与多于一个基因的转录物比对。我们决定检查注释的所有常染色体和X染色体系统地使用生物信息学methods.Results:出42683记者,我们发现,25505(60%)通过了我们所有的测试,被认为是“完全有效”。9964(23%)报告基因没有有意义的标识符,定位到错误的染色体,或没有通过基本的比对测试,阻止我们将这些报告基因的表达值与唯一注释的人类基因相关联。剩余的7214个(17%)报告基因可能与一个独特的基因或一个独特的基因间位置相关,但无法映射到RefSeq中的转录本。7214名记者进一步划分为三个不同水平的validity.Conclusion:表达阵列研究应评估的注释的报告,并删除那些有可疑的注释的报告。这种评估可以系统地或半自动地进行,但必须认识到,数据源经常更新,导致验证结果随时间略有变化。
Background: While attempting to reanalyze published data from Agilent 4 x 44 human expression chips, we found that some of the 60-mer olignucleotide features could not be interpreted as representing single human genes. For example, some of the oligonucleotides align with the transcripts of more than one gene. We decided to check the annotations for all autosomes and the X chromosome systematically using bioinformatics methods.Results: Out of 42683 reporters, we found that 25505 (60%) passed all our tests and are considered "fully valid". 9964 (23%) reporters did not have a meaningful identifier, mapped to the wrong chromosome, or did not pass basic alignment tests preventing us from correlating the expression values of these reporters with a unique annotated human gene. The remaining 7214 (17%) reporters could be associated with either a unique gene or a unique intergenic location, but could not be mapped to a transcript in RefSeq. The 7214 reporters are further partitioned into three different levels of validity.Conclusion: Expression array studies should evaluate the annotations of reporters and remove those reporters that have suspect annotations. This evaluation can be done systematically and semiautomatically, but one must recognize that data sources are frequently updated leading to slightly changing validation results over time.