Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts.
Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts.
复制标题
DOI:
10.1186/1471-2105-6-103
复制
发表时间:
2005-04-22
影响因子:
3
通讯作者:
Spackman K
中科院分区:
文献类型:
--
作者:
Cohen AM;Hersh WR;Dubay C;Spackman K
Text-mining can assist biomedical researchers in reducing information overload by extracting useful knowledge from large collections of text. We developed a novel text-mining method based on analyzing the network structure created by symbol co-occurrences as a way to extend the capabilities of knowledge extraction. The method was applied to the task of automatic gene and protein name synonym extraction. Performance was measured on a test set consisting of about 50,000 abstracts from one year of MEDLINE. Synonyms retrieved from curated genomics databases were used as a gold standard. The system obtained a maximum F-score of 22.21% (23.18% precision and 21.36% recall), with high efficiency in the use of seed pairs. The method performs comparably with other studied methods, does not rely on sophisticated named-entity recognition, and requires little initial seed knowledge.
登录
查看更多内容
影响因子:
5.3
作者:
Povey, S;Lovering, R;Wain, H
通讯作者:
Wain, H
影响因子:
5.8
作者:
Kolpakov, FA;Ananko, EA;Kolchanov, NA
通讯作者:
Kolchanov, NA
影响因子:
4.5
作者:
Hirschman, L;Morgan, AA;Yeh, AS
通讯作者:
Yeh, AS
影响因子:
14.9
作者:
Gelbart, W;Bayraktaroglu, L;Wiel, C
通讯作者:
Wiel, C
影响因子:
14.9
作者:
Kanehisa, M;Goto, S
通讯作者:
Goto, S