Free Text Gene Name Recognition
Free Text Gene Name Recognition
批准号:
7594470
负责人:
willy john wilbur
金额:
$19.42万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
BiologyCategoriesClassDatabasesEducational workshopGene ProteinsGenesGeneticHarvestInternetLanguageLearningLiteratureMEDLINEMaterials TestingMethodsModelingNamesNatural Language ProcessingNumbersPerformancePhaseProteinsProteomicsPurposeRangeSamplingScoreSemanticsSiteSpainSystemTechniquesTextThinkingTrainingVariantWorkabstractingbasedesignimprovedmarkov modelresearch study
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Currently we are pursuing two projects designed to make progress on the problem of gene/protein name recognition:
1) We have produced a set of 20,000 sentences with all occurrences of gene/protein names in them marked up with the character offset for name beginning and name ending in the sentence. The sentences were taken as random samples from restricted classes of MEDLINE abstracts. Half were chosen as likely to have gene/protein names in them and half were selected as unlikely to have such names. Since there is ambiguity in marking names, alternative markings are listed as correct answers where this is thought to be appropriate. Three fourths of these names formed the basis for a task in the BioCreAtIvE1 (Critical Assessment of Information Extraction in Biology) Workshop held in Granada, Spain in 2004. Twelve teams attempted to designed systems that could correctly tag the gene/protein names in the sentences. Several teams obtained precisions and recalls in the low 80% range. A number of different approaches were successful and these results suggest ways in which gene/protein name tagging. The 20,000 sentences forming the basis of this work have been re-edited and a number of errors corrected. The 15,000 sentences which formed the basis of BioCreAtIvE1 and currently being used for the training phase of BioCreAtIvE2 and the last 5,000 sentences which have never been released will form the testing material for the BioCreAtIvE2 which is planned for early 2007.
2) We have become convinced that more information about the different types of entities that can occur in sentences in MEDLINE can be used to improve name recognition. This has led us to design a set of semantic categories and to attempt to fill these categories with actual names that can be harvested from databases and from web sites. We call the result SEMCAT. It currently recognizes seventy-five categories and contains about five million name strings distributed over those categories. We have experimented with probabilistic context free grammars and Markov models of text strings in an attempt to learn how to recognize the entities in different categories. In order to improve performance we have developed a new model term a Priority Model for name recognition. This model allows us to categorize names as gene/protein names with an F-score of 0.96 and better then what we were able to achieve with either a language model of a probabilistic context free grammar. We are currently using this to create features for using in a conditional random fields approach to gene/protein name recognition and are achieving about an 0.83 F-score.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A DOCUMENT PROCESSING SYSTEM
-
批准号:6111059
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6111081
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6290494
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical phrase extraction techniques in natural language databases.
-
批准号:6228047
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6546806
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme recognition in document collections.
-
批准号:6432765
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:6843586
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:7316268
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7148024
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7148023
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7594456
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6843558
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6843556
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6432759
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6546807
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
AUTOMATIC BAYESIAN METHODS IN TEXT RETRIEVAL
-
批准号:6432747
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7316262
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:7316265
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7148039
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections.
-
批准号:7594468
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
海外基金