Mistaken identifiers: Gene name errors can be introduced inadvertently when using Excel in bioinformatics

Mistaken identifiers: Gene name errors can be introduced inadvertently when using Excel in bioinformatics
复制标题

DOI:
10.1186/1471-2105-5-80
复制
发表时间:
2004-06-23
期刊:
影响因子:
3
通讯作者:
Weinstein, JN
Weinstein, JN
中科院分区:
生物学4区
文献类型:
--
作者:
Zeeberg, BR;Riss, J;Weinstein, JN

文献摘要

被引文献

相似文献

背景资料:当处理微阵列数据集,我们最近注意到,一些基因名称被无意中更改为non-gene names.Results:一个小侦探工作跟踪问题的默认日期格式转换和浮点格式转换在非常有用的Excel程序包。日期转换至少影响30个基因名称;如果包括Riken标识符,浮点转换至少影响2,000个。这些转换是不可逆的,原始的基因名称不能recovery.Conclusions:Excel的用户分析涉及基因名称应该知道这个问题,这可能会导致基因,包括医学上重要的,从视图中丢失,甚至污染精心策划的公共数据库。我们提供了解决问题的方法和脚本。
Background: When processing microarray data sets, we recently noticed that some gene names were being changed inadvertently to non-gene names.Results: A little detective work traced the problem to default date format conversions and floating-point format conversions in the very useful Excel program package. The date conversions affect at least 30 gene names; the floating-point conversions affect at least 2,000 if Riken identifiers are included. These conversions are irreversible; the original gene names cannot be recovered.Conclusions: Users of Excel for analyses involving gene names should be aware of this problem, which can cause genes, including medically important ones, to be lost from view and which has contaminated even carefully curated public databases. We provide work-arounds and scripts for circumventing the problem.