Free Text Gene Name Recognition
Free Text Gene Name Recognition
批准号:
8558107
负责人:
Willy Wilbur
金额:
$19.52万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AlgorithmsBiological SciencesBiologyCommunitiesDatabasesDependencyDevelopmentDiseaseEducational workshopElementsEvaluationGenbankGene ProteinsGenesGeneticGoalsGoldHumanJudgmentLiteratureMachine LearningMapsMeasuresMethodsModelingNamesNatural Language ProcessingPaperParticipantPerformanceProcessProteinsProteomicsPubMedRecordsResearch PersonnelSemanticsSystemTechniquesTextTriageVariantWorkWritingabstractingbaseimprovedinnovationprotein protein interactionsymposiumtext searchingtool
中文摘要
1)自2005年以来,我一直是BioCreative研讨会的联合组织者,并参加了BioCreative II(2007)、BioCreative III(2010)和BioCreative-2012研讨会(2012)。生物创意讲习班的总体目标是促进开发对生物科学研究人员和数据库馆长社区有用的文本挖掘和文本处理工具。我们对BioCreative III的贡献是组织了基因标准化(GN)任务。这包括选择要添加注释的文件、监督注释过程、评估与会者提交的材料、编写任务的完整说明并在会议上介绍结果。我们在这项任务中引入了两项创新。首先,我们使用John Spouge和他的团队开发的Tap-k测量,并引入了一种EM算法来根据任务的所有参与者条目估计正确答案。Tap-k是一种性能度量,它可以最好地描述为截断的平均平均精度,以及用于检索上次有用命中以下的无用记录的惩罚项。EM算法允许我们在一组比我们能为其提供黄金标准人类判断的全文文档上评估人们的预测(507)。我们还提供了一组50个具有人工注释的全文文档。当EM算法评估的结果与金标准结果进行比较时,排名相当接近,这使我们可以得出结论,自动评估方法成功地挑出了性能最好的系统。我们还进入了分类任务,以检测适合管理蛋白质-蛋白质相互作用的论文。对于这项任务,我们使用优先级模型来识别基因/蛋白质名称,并使用句法分析来准备蛋白质与其他文本元素之间的依存关系,这些关系以及文本单词被用作特征。然后将机器学习应用到这个表示中,我们提交了任务中的最佳表现。我们最近帮助组织了与2012年生物认证大会相关的BioCreative-2012研讨会,并参与了任务I,该任务涉及为CTD数据库制作一个分类系统。我们的方法是基于BioCreative III的GN任务所使用的方法,但我们采用了一种不同的方法来基于语义分类器识别基因、蛋白质和疾病,并且我们还基于CTD数据库的LDA分析添加了特征。我们的方法在获得任务的最佳分类结果方面是有效的。
2)我们最近联合主持了BioCreative III研讨会,在该研讨会中,主要竞争任务是在全文文章中查找提到基因的内容,并将其映射到其GenBank标识符并根据可靠性对其进行评分,将PubMed记录归类为可能代表包含蛋白质-蛋白质相互作用信息的文章,以及在全文中找到描述实验人员用来验证蛋白质-蛋白质相互作用的方法的文本。我们组织了第一项任务,并参与了第二项任务。在第二个任务中,我们使用优先级模型来定位蛋白质提及,它被证明是非常成功的,与其他方法相比具有竞争力。
3)我们目前正在努力开发更通用的方法,根据摘要为PPI寻找高价值文章。这一努力不仅涉及更强大的排名方法,还包括向用户显示证据以供用户快速评估的方式。
英文摘要
1) I have been a co-organizer of the BioCreative Workshops since 2005 and have taken part in BioCreative II (2007), BioCreative III (2010), and the BioCreative-2012 Workshop (2012). The overall goal of the BioCreative Workshops is to promote the development of text mining and text processing tools which are useful to the communities of researchers and database curators in the biological sciences. Our contribution for BioCreative III was to organize the Gene Normalization (GN) task. This included selecting documents to annotate, overseeing the annotation process, evaluating participants submissions, and writing up a full description of the task and presenting results at the conference. We introduced two innovations in the task. First, we used the Tap-k measure developed by John Spouge and his group and we introduced an EM algorithm to estimate the correct answers based on all participant entries for the task. Tap-k is a performance measure which can best be characterized as a truncated mean average precision with a penalty term for retrieving useless records below the last useful hit. The EM algorithm allowed us to evaluate peoples predictions over a much larger set of full text documents than we could provide gold standard human judgments for (507). We also provided a set of 50 full text documents for which we had human annotations. When the results of the EM algorithm evaluation were compared with the gold standard results the ranks were quite close and allowed us to conclude that the automatic method of assessment was successful in singling out the top performing systems. We also entered the triage task for detecting papers suitable for curation of protein-protein interactions. For this task we used the priority model to identify gene/protein names and used parsing to prepare dependency relations between proteins and other text elements and these relationships as well as text words were used as features. Machine learning was then applied to this representation and we turned in the best performance on the task. We recently helped organize the BioCreative-2012 Workshop associated with the Biocuration 2012 Conference and also participated in Task I which involved producing a triage system for the CTD database. Our approach was based on the approach used for the GN task of BioCreative III, but we took a different approach to identify genes, proteins, and diseases based on a semantic classifier and we also added features based on an LDA analysis of the CTD database. Our approach was effective in obtaining the best triage results on the task.
2) We recently co-chaired the BioCreative III Workshop in which the main competitive tasks were to find gene mentions in a full text article and map them to their GenBank identifiers and score them as to reliability, to classify PubMed records as likely to represent articles containing information on protein-protein interactions, and to find the text in full papers that describes the method used by an experimenter to experimentally verify a protein-protein interaction. We organized the first of these task and participated in the second. In the second task we used the priority model to locate protein mentions and it proved very successful and competitive with other approaches.
3) We are currently working to develop more general methods of finding high value articles for PPI based on their abstracts. This effort involves not only more powerful ranking methods, but also ways to display evidence to the user for a users quick evaluation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8558105
-
项目类别:
-
资助金额:$56.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Natural Language Processing Techniques To Enhance Information Access.
-
批准号:8943224
-
项目类别:
-
资助金额:$56.15万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8344960
-
项目类别:
-
资助金额:$23.98万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8344939
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7969244
-
项目类别:
-
资助金额:$77.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8149591
-
项目类别:
-
资助金额:$13.71万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8149592
-
项目类别:
-
资助金额:$17.63万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8149602
-
项目类别:
-
资助金额:$47.01万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:9160906
-
项目类别:
-
资助金额:$43.97万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:7969199
-
项目类别:
-
资助金额:$20.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8344948
-
项目类别:
-
资助金额:$59.96万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:8344950
-
项目类别:
-
资助金额:$17.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8558117
-
项目类别:
-
资助金额:$26.03万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8943215
-
项目类别:
-
资助金额:$18.72万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7969197
-
项目类别:
-
资助金额:$12.9万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8344938
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:9160916
-
项目类别:
-
资助金额:$14.32万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:9160928
-
项目类别:
-
资助金额:$12.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7735088
-
项目类别:
-
资助金额:$29.31万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Enhancement
-
批准号:8177730
-
项目类别:
-
资助金额:$48.97万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
海外基金