Multi-relational learning, text mining, and semi-supervised learning for functional genomics

Multi-relational learning, text mining, and semi-supervised learning for functional genomics
复制标题

DOI:
10.1023/b:mach.0000035472.73496.0c
复制
发表时间:
2004-10-01
期刊:
影响因子:
7.5
通讯作者:
Scheffer, T
Scheffer, T
中科院分区:
计算机科学3区
文献类型:
--
作者:
Krogel, MA;Scheffer, T

文献摘要

被引文献

相似文献

我们专注于预测与酵母基因组中的基因相对应的蛋白质的功能特性的问题。我们的目标是研究利用该问题设置中可用的所有数据源的方法的有效性,包括关系数据、研究论文摘要和未标记数据。我们研究了一种使用关系基因相互作用数据的命题化方法。我们研究文本分类和信息提取对于利用科学摘要集合的好处。我们研究使用未标记数据的转导和协同训练。我们报告了所调查方法的积极和消极结果。研究的任务是 2001 年和 2002 年的 KDD Cup 任务。我们描述的解决方案在 2001 年获得了任务 2 的最高分,2001 年任务 3 获得了第四名,两个子任务之一获得了最高分,2002 年整个任务 2 获得了第三名。
We focus on the problem of predicting functional properties of the proteins corresponding to genes in the yeast genome. Our goal is to study the effectiveness of approaches that utilize all data sources that are available in this problem setting, including relational data, abstracts of research papers, and unlabeled data. We investigate a propositionalization approach which uses relational gene interaction data. We study the benefit of text classification and information extraction for utilizing a collection of scientific abstracts. We study transduction and co-training for using unlabeled data. We report on both, positive and negative results on the investigated approaches. The studied tasks are KDD Cup tasks of 2001 and 2002. The solutions which we describe achieved the highest score for task 2 in 2001, the fourth rank for task 3 in 2001, the highest score for one of the two subtasks and the third place for the overall task 2 in 2002.