Transactional database transformation and its application in prioritizing human disease genes.

Transactional database transformation and its application in prioritizing human disease genes.
复制标题

DOI:
10.1109/tcbb.2011.58
复制
发表时间:
2012-01
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
Huang K
Huang K
中科院分区:
其他
文献类型:
--
作者:
Xiang Y;Payne PR;Huang K

文献摘要

相似文献

通常称为事务数据库的二进制(0,1)矩阵可以表示许多应用数据,包括基因-表型数据,其中“1”表示已确认的基因-表型关系,而“0”表示未知关系。人们很自然地会问,在这些“0”S和“1”S背后隐藏着什么信息。遗憾的是,最近的矩阵补全方法虽然在许多情况下非常有效,但不太可能从这些(0,1)-矩阵中推断出一些有趣的东西。为了应对这一挑战,我们提出了一种非常简洁和有效的算法IndEvi来执行基于独立证据的事务数据库转换。(0,1)-矩阵的每个条目通过从该条目的整个矩阵中提取的“独立证据”(最大支持模式)来评估。条目的值,无论其值是0还是1,对其独立证据都完全没有影响。在基因表型数据库上的实验表明,该方法在候选基因排序和未知疾病基因预测方面具有很高的应用前景。
Binary (0,1) matrices, commonly known as transactional databases, can represent many application data, including gene-phenotype data where “1” represents a confirmed gene-phenotype relation and “0” represents an unknown relation. It is natural to ask what information is hidden behind these “0”s and “1”s. Unfortunately, recent matrix completion methods, though very effective in many cases, are less likely to infer something interesting from these (0,1)-matrices. To answer this challenge, we propose IndEvi, a very succinct and effective algorithm to perform independent-evidence-based transactional database transformation. Each entry of a (0,1)-matrix is evaluated by “independent evidence” (maximal supporting patterns) extracted from the whole matrix for this entry. The value of an entry, regardless of its value as 0 or 1, has completely no effect for its independent evidence. The experiment on a gene-phenotype database shows that our method is highly promising in ranking candidate genes and predicting unknown disease genes.