Challenges in the association of human single nucleotide polymorphism mentions with unique database identifiers.

Challenges in the association of human single nucleotide polymorphism mentions with unique database identifiers.
复制标题

DOI:
10.1186/1471-2105-12-s4-s4
复制
发表时间:
2011
期刊:
影响因子:
3
通讯作者:
Friedrich CM
Friedrich CM
中科院分区:
生物学4区
文献类型:
--
作者:
Thomas PE;Klinger R;Furlong LI;Hofmann-Apitius M;Friedrich CM

文献摘要

被引文献

相似文献

大多数关于基因组变异及其与表型的关联的信息仅在科学出版物中涵盖,而不是在结构化数据库中。这些文本通常使用自然语言描述变化;数据库标识符很少被提及。这使变异、相关文章的检索以及信息提取(例如寻找生物学含义)变得复杂。为了克服这些挑战,需要开发将变体的文本提及映射到数据库标识符的过程。本文描述了一种变体提及规范化的工作流程,即它们与唯一数据库标识符的关联。常见的陷阱在解释单核苷酸多态性(SNP)提到强调和讨论。在包含527个snp提及的296篇MEDLINE摘要的文本语料库中,开发的规范化程序实现了与dbSNP标识符明确关联的精度为98.1%,召回率为67.5%。带注释的语料库可在http://www.scai.fraunhofer.de/snp-normalization-corpus.html免费获得。类似的方法通常侧重于蛋白质序列上提到的变异,而忽略了其他SNP提到的问题。本文的结果表明,在DNA水平上描述的snp的正常化比在蛋白质水平上描述的snp的正常化更困难。规范化的挑战体现在语料库中出现的歧义和错误。
Most information on genomic variations and their associations with phenotypes are covered exclusively in scientific publications rather than in structured databases. These texts commonly describe variations using natural language; database identifiers are seldom mentioned. This complicates the retrieval of variations, associated articles, as well as information extraction, e. g. the search for biological implications. To overcome these challenges, procedures to map textual mentions of variations to database identifiers need to be developed. This article describes a workflow for normalization of variation mentions, i.e. the association of them to unique database identifiers. Common pitfalls in the interpretation of single nucleotide polymorphism (SNP) mentions are highlighted and discussed. The developed normalization procedure achieves a precision of 98.1 % and a recall of 67.5% for unambiguous association of variation mentions with dbSNP identifiers on a text corpus based on 296 MEDLINE abstracts containing 527 mentions of SNPs. The annotated corpus is freely available at http://www.scai.fraunhofer.de/snp-normalization-corpus.html. Comparable approaches usually focus on variations mentioned on the protein sequence and neglect problems for other SNP mentions. The results presented here indicate that normalizing SNPs described on DNA level is more difficult than the normalization of SNPs described on protein level. The challenges associated with normalization are exemplified with ambiguities and errors, which occur in this corpus.