BioC: a minimalist approach to interoperability for biomedical text processing.

BioC: a minimalist approach to interoperability for biomedical text processing.
复制标题

DOI:
10.1093/database/bat064
复制
发表时间:
2013
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Wilbur WJ
Wilbur WJ
中科院分区:
其他
文献类型:
--
作者:
Comeau DC;Islamaj Doğan R;Ciccarese P;Cohen KB;Krallinger M;Leitner F;Lu Z;Peng Y;Rinaldi F;Torii M;Valencia A;Verspoor K;Wiegers TC;Wu CH;Wilbur WJ

文献摘要

参考文献

被引文献

相似文献

大量的科学信息被编码在自然语言文本中,并且这样的文本的数量已经变得如此之大,以至于在搜索过程中将人类作为第一步在经济上不再可行。自然语言处理和文本挖掘工具对于促进从文本中搜索和提取信息至关重要。这导致了积极的研究努力,以创建有用的工具,并创建人性化标记的文本语料库,这可以用来改善这些工具。为了鼓励将这些努力组合成更大、更强大和更有能力的系统,非常需要一种通用的交换格式,以便以简单的方式在不同的语言处理系统和文本挖掘工具之间表示、存储和交换数据。在这里,我们提出了一个简单的可扩展标记语言格式来共享文本文档和注释。所提出的注释方法允许表示大量不同的注释,包括句子、标记、词性、命名实体(诸如基因或疾病)以及命名实体之间的关系。此外,我们还提供了简单的代码来保存这些数据,从可扩展标记语言文件中读取数据并将其写回可扩展标记语言文件,以及执行一些示例处理。我们还描述了已完成的以及正在进行的工作,在几个方向应用的方法。代码和数据可在http://bioc.sourceforge.net/上获得。数据库URL:http://bioc.sourceforge.net/
A vast amount of scientific information is encoded in natural language text, and the quantity of such text has become so great that it is no longer economically feasible to have a human as the first step in the search process. Natural language processing and text mining tools have become essential to facilitate the search for and extraction of information from text. This has led to vigorous research efforts to create useful tools and to create humanly labeled text corpora, which can be used to improve such tools. To encourage combining these efforts into larger, more powerful and more capable systems, a common interchange format to represent, store and exchange the data in a simple manner between different language processing systems and text mining tools is highly desirable. Here we propose a simple extensible mark-up language format to share text documents and annotations. The proposed annotation approach allows a large number of different annotations to be represented including sentences, tokens, parts of speech, named entities such as genes or diseases and relationships between named entities. In addition, we provide simple code to hold this data, read it from and write it back to extensible mark-up language files and perform some sample processing. We also describe completed as well as ongoing work to apply the approach in several directions. Code and data are available at http://bioc.sourceforge.net/. Database URL: http://bioc.sourceforge.net/
生物公约概述:生物学信息提取的批判性评估。
DOI: 10.1186/1471-2105-6-s1-s1
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A
通讯作者: Valencia A
DOI: 10.1093/nar/gks994
发表时间: 2013-01
影响因子: 14.9
作者:
Davis AP;Murphy CG;Johnson R;Lay JM;Lennon-Hopkins K;Saraceni-Richards C;Sciaky D;King BL;Rosenstein MC;Wiegers TC;Mattingly CJ
通讯作者: Mattingly CJ
DOI: 10.1186/2041-1480-3-s1-s1
发表时间: 2012-04-24
影响因子: 1.9
作者:
Ciccarese P;Ocana M;Clark T
通讯作者: Clark T
重构语料库:一项可行性研究。
DOI: 10.1186/1747-5333-2-4
发表时间: 2007-09-13
期刊: Journal of biomedical discovery and collaboration
影响因子: --
作者:
Johnson HL;Baumgartner WA Jr;Krallinger M;Cohen KB;Hunter L
通讯作者: Hunter L
DOI: 10.1186/2041-1480-2-s2-s4
发表时间: 2011-05-17
影响因子: 1.9
作者:
Ciccarese P;Ocana M;Garcia Castro LJ;Das S;Clark T
通讯作者: Clark T