Free Text Gene Name Recognition
Free Text Gene Name Recognition
批准号:
7316268
负责人:
willy john wilbur
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
目前,我们正在进行两个旨在解决基因/蛋白质名称识别问题的项目:
1)我们已经生成了一组20,000个句子,其中所有出现的基因/蛋白质名称都用句子中名称开头和名称结尾的字符偏移标记。这些句子是从MEDLINE摘要的限制类中随机抽取的。其中一半被选为可能有基因/蛋白质名称,另一半被选为不太可能有这样的名称。由于在标记名称时存在歧义,因此在认为适当的情况下,将替代标记列为正确答案。这些名称的四分之三构成了2004年在西班牙格拉纳达举行的BioCreAtIvE 1(生物学信息提取的批判性评估)讲习班的一项任务的基础。12个团队试图设计能够正确标记句子中基因/蛋白质名称的系统。几个团队获得了80%以下的精确度和召回率。许多不同的方法是成功的,这些结果表明,基因/蛋白质名称标记的方式。构成本书基础的20,000个句子已重新编辑,并纠正了一些错误。构成BioCreAtIvE 1基础的15,000个句子目前正用于BioCreAtIvE 2的培训阶段,而最后5,000个从未发布的句子将构成计划于2007年初发布的BioCreAtIvE 2的测试材料。
2)我们已经确信,可以使用MEDLINE中句子中可能出现的不同类型实体的更多信息来提高名称识别。这使我们设计了一组语义类别,并试图用可以从数据库和网站中获得的实际名称来填充这些类别。我们称之为SEMCAT。它目前可以识别75个类别,包含分布在这些类别中的大约500万个名称字符串。我们尝试了概率上下文无关文法和文本字符串的马尔可夫模型,试图学习如何识别不同类别的实体。为了提高性能,我们开发了一个新的模型术语优先级模型的名称识别。该模型允许我们将名称分类为基因/蛋白质名称,F分数为0.96,并且比我们能够用概率上下文无关语法的语言模型实现的更好。我们目前正在使用它来创建用于基因/蛋白质名称识别的条件随机场方法的功能,并实现了约0.83的F分数。
英文摘要
Currently we are pursuing two projects designed to make progress on the problem of gene/protein name recognition:
1) We have produced a set of 20,000 sentences with all occurrences of gene/protein names in them marked up with the character offset for name beginning and name ending in the sentence. The sentences were taken as random samples from restricted classes of MEDLINE abstracts. Half were chosen as likely to have gene/protein names in them and half were selected as unlikely to have such names. Since there is ambiguity in marking names, alternative markings are listed as correct answers where this is thought to be appropriate. Three fourths of these names formed the basis for a task in the BioCreAtIvE1 (Critical Assessment of Information Extraction in Biology) Workshop held in Granada, Spain in 2004. Twelve teams attempted to designed systems that could correctly tag the gene/protein names in the sentences. Several teams obtained precisions and recalls in the low 80% range. A number of different approaches were successful and these results suggest ways in which gene/protein name tagging. The 20,000 sentences forming the basis of this work have been re-edited and a number of errors corrected. The 15,000 sentences which formed the basis of BioCreAtIvE1 and currently being used for the training phase of BioCreAtIvE2 and the last 5,000 sentences which have never been released will form the testing material for the BioCreAtIvE2 which is planned for early 2007.
2) We have become convinced that more information about the different types of entities that can occur in sentences in MEDLINE can be used to improve name recognition. This has led us to design a set of semantic categories and to attempt to fill these categories with actual names that can be harvested from databases and from web sites. We call the result SEMCAT. It currently recognizes seventy-five categories and contains about five million name strings distributed over those categories. We have experimented with probabilistic context free grammars and Markov models of text strings in an attempt to learn how to recognize the entities in different categories. In order to improve performance we have developed a new model term a Priority Model for name recognition. This model allows us to categorize names as gene/protein names with an F-score of 0.96 and better then what we were able to achieve with either a language model of a probabilistic context free grammar. We are currently using this to create features for using in a conditional random fields approach to gene/protein name recognition and are achieving about an 0.83 F-score.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A DOCUMENT PROCESSING SYSTEM
-
批准号:6111059
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6111081
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6290494
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical phrase extraction techniques in natural language databases.
-
批准号:6228047
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6546806
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme recognition in document collections.
-
批准号:6432765
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:6843586
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7148024
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7148023
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7594456
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7316262
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:7316265
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
AUTOMATIC BAYESIAN METHODS IN TEXT RETRIEVAL
-
批准号:6432747
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6432759
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6546807
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6843556
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6843558
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7148039
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections.
-
批准号:7594468
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A DOCUMENT PROCESSING SYSTEM
-
批准号:6290480
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
国内基金
海外基金
登录
查看更多内容
J-TEXT托卡马克上边界湍流与撕裂模相互作用的实验研究
-
批准号:12375223
-
项目类别:面上项目
-
资助金额:54万元
-
批准年份:2023
-
负责人:刘海
-
依托单位:
J-TEXT装置外加三维磁场主动调控偏滤器脱靶的实验研究
-
批准号:12305243
-
项目类别:青年科学基金项目
-
资助金额:20万元
-
批准年份:2023
-
负责人:周松
-
依托单位:
J-TEXT托卡马克装置上多模式磁扰动对逃逸电流影响研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:林志芳
-
依托单位:
J-TEXT托卡马克上边界湍流特性对高密度运行影响的实验研究
-
批准号:11905080
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2019
-
负责人:石鹏
-
依托单位:
关于J-TEXT托卡马克上微撕裂模电磁湍流及其输运的实验研究
-
批准号:11605067
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2016
-
负责人:陈杰
-
依托单位:
基于J-TEXT远红外偏振干涉仪的相干散射与密度扰动的实验研究
-
批准号:11575067
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:高丽
-
依托单位:
J-TEXT上外加磁扰动抑制等离子体破裂下逃逸电子产生的实验研究
-
批准号:11275079
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:陈忠勇
-
依托单位:
J-TEXT托卡马克等离子体粒子输运的密度调制实验研究
-
批准号:11105056
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2011
-
负责人:高丽
-
依托单位: