Annotating genes and genomes with DNA sequences extracted from biomedical articles.

Annotating genes and genomes with DNA sequences extracted from biomedical articles.
复制标题

DOI:
10.1093/bioinformatics/btr043
复制
发表时间:
2011-04-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Bergman CM
Bergman CM
中科院分区:
其他
文献类型:
--
作者:
Haeussler M;Gerner M;Bergman CM

文献摘要

参考文献

被引文献

相似文献

动机:出版物和DNA测序的增加使得为特定基因或基因组区域找到相关文章的问题比以往任何时候都更具挑战性。现有的文本挖掘方法主要集中在寻找英文文本中的基因名称或标识符。这些通常不是唯一的,并且不能识别研究的确切基因组位置。结果如下:在这里,我们报告了一种新的文本挖掘方法的结果,该方法从生物医学文章中提取DNA序列,并将其自动映射到基因组数据库。我们发现,PubMed central(PMC)中大约20%的开放获取文章具有可提取的DNA序列,这些序列可以准确地映射到正确的基因(91%)和基因组(96%)。我们说明了text 2genome从超过150 000 PMC文章中提取的数据用于解释ChIP-seq数据和设计定量逆转录酶(RT)-PCR实验的实用性。结论:我们的方法将文章链接到基因和生物体,而不依赖于基因名称或标识符。它还生成生物医学文献的基因组注释轨迹,从而使研究人员能够使用现代基因组浏览器的功能来访问和分析基因组数据背景下的出版物。可用性和实施:源代码可在BSD许可证下从http://sourceforge.net/projects/text2genome/获得,结果可在www.example.com浏览和下载。联系方式:maximilianh@gmail.com补充信息:补充数据可在生物信息学在线获得。http://text2genome.org
Motivation: Increasing rates of publication and DNA sequencing make the problem of finding relevant articles for a particular gene or genomic region more challenging than ever. Existing text-mining approaches focus on finding gene names or identifiers in English text. These are often not unique and do not identify the exact genomic location of a study. Results: Here, we report the results of a novel text-mining approach that extracts DNA sequences from biomedical articles and automatically maps them to genomic databases. We find that ∼20% of open access articles in PubMed central (PMC) have extractable DNA sequences that can be accurately mapped to the correct gene (91%) and genome (96%). We illustrate the utility of data extracted by text2genome from more than 150 000 PMC articles for the interpretation of ChIP-seq data and the design of quantitative reverse transcriptase (RT)-PCR experiments. Conclusion: Our approach links articles to genes and organisms without relying on gene names or identifiers. It also produces genome annotation tracks of the biomedical literature, thereby allowing researchers to use the power of modern genome browsers to access and analyze publications in the context of genomic data. Availability and implementation: Source code is available under a BSD license from http://sourceforge.net/projects/text2genome/ and results can be browsed and downloaded at http://text2genome.org. Contact: maximilianh@gmail.com Supplementary information: Supplementary data are available at Bioinformatics online.
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1101/gr.6.10.995
发表时间: 1996-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Gibson, UEM;Heid, CA;Williams, PM
通讯作者: Williams, PM
UCSC基因组浏览器数据库:更新2010。
DOI: 10.1093/nar/gkp939
发表时间: 2010-01
影响因子: 14.9
作者:
Rhead B;Karolchik D;Kuhn RM;Hinrichs AS;Zweig AS;Fujita PA;Diekhans M;Smith KE;Rosenbloom KR;Raney BJ;Pohl A;Pheasant M;Meyer LR;Learned K;Hsu F;Hillman-Jackson J;Harte RA;Giardine B;Dreszer TR;Clawson H;Barber GP;Haussler D;Kent WJ
通讯作者: Kent WJ
DOI: 10.1186/1471-2105-6-s1-s12
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Colosimo ME;Morgan AA;Yeh AS;Colombe JB;Hirschman L
通讯作者: Hirschman L
DOI: 10.1093/nar/gkp871
发表时间: 2010-01
影响因子: 14.9
作者:
Kersey PJ;Lawson D;Birney E;Derwent PS;Haimel M;Herrero J;Keenan S;Kerhornou A;Koscielny G;Kähäri A;Kinsella RJ;Kulesha E;Maheswari U;Megy K;Nuhn M;Proctor G;Staines D;Valentin F;Vilella AJ;Yates A
通讯作者: Yates A