Free Text Gene Name Recognition
Free Text Gene Name Recognition
批准号:
7594470
负责人:
willy john wilbur
金额:
$19.42万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
BiologyCategoriesClassDatabasesEducational workshopGene ProteinsGenesGeneticHarvestInternetLanguageLearningLiteratureMEDLINEMaterials TestingMethodsModelingNamesNatural Language ProcessingNumbersPerformancePhaseProteinsProteomicsPurposeRangeSamplingScoreSemanticsSiteSpainSystemTechniquesTextThinkingTrainingVariantWorkabstractingbasedesignimprovedmarkov modelresearch study
中文摘要
目前,我们正在进行两个项目,旨在在基因/蛋白质名称识别问题上取得进展:
1)我们制作了一组20,000个句子,其中所有出现的基因/蛋白质名称都在句子中标记了名称开头和结尾的字符偏移量。这些句子是从限制类别的MEDLINE摘要中随机抽取的。其中一半被选为可能有基因/蛋白质名称,另一半被选为不太可能有这样的名称。由于在标记名称时存在歧义,因此在认为合适的情况下,可选标记被列为正确答案。其中四分之三的名字构成了2004年在西班牙格拉纳达举行的BioCreAtIVE1(生物学信息提取关键评估)讲习班的一项任务的基础。12个团队试图设计出能够正确标记句子中基因/蛋白质名称的系统。有几个团队的准确率和召回率都在80%以下。许多不同的方法都取得了成功,这些结果表明了基因/蛋白质名称标签的方法。构成这项工作基础的20,000个句子已重新编辑,并更正了一些错误。这15,000个句子构成了BioCreAtIVE1的基础,目前正用于BioCreAtIVE2的训练阶段,最后5,000个句子从未发布,将构成计划于2007年初举行的BioCreAtIVE2的测试材料。
2)我们相信,关于MEDLINE中句子中可能出现的不同类型实体的更多信息可以用于改善姓名识别。这导致我们设计了一套语义类别,并试图用可以从数据库和网站获得的实际名称填充这些类别。我们称结果为SEMCAT。它目前识别75个类别,包含分布在这些类别中的大约500万个名称字符串。我们已经使用概率上下文无关文法和文本串的马尔可夫模型进行了实验,试图学习如何识别不同类别的实体。为了提高性能,我们开发了一个新的模型术语,即用于姓名识别的优先模型。这个模型允许我们将名字归类为F分数为0.96的基因/蛋白质名字,比我们使用概率上下文无关文法的语言模型所能达到的效果更好。我们目前正在使用它来创建在条件随机场方法中用于基因/蛋白质名称识别的特征,并正在获得大约0.83F-Score。
英文摘要
Currently we are pursuing two projects designed to make progress on the problem of gene/protein name recognition:
1) We have produced a set of 20,000 sentences with all occurrences of gene/protein names in them marked up with the character offset for name beginning and name ending in the sentence. The sentences were taken as random samples from restricted classes of MEDLINE abstracts. Half were chosen as likely to have gene/protein names in them and half were selected as unlikely to have such names. Since there is ambiguity in marking names, alternative markings are listed as correct answers where this is thought to be appropriate. Three fourths of these names formed the basis for a task in the BioCreAtIvE1 (Critical Assessment of Information Extraction in Biology) Workshop held in Granada, Spain in 2004. Twelve teams attempted to designed systems that could correctly tag the gene/protein names in the sentences. Several teams obtained precisions and recalls in the low 80% range. A number of different approaches were successful and these results suggest ways in which gene/protein name tagging. The 20,000 sentences forming the basis of this work have been re-edited and a number of errors corrected. The 15,000 sentences which formed the basis of BioCreAtIvE1 and currently being used for the training phase of BioCreAtIvE2 and the last 5,000 sentences which have never been released will form the testing material for the BioCreAtIvE2 which is planned for early 2007.
2) We have become convinced that more information about the different types of entities that can occur in sentences in MEDLINE can be used to improve name recognition. This has led us to design a set of semantic categories and to attempt to fill these categories with actual names that can be harvested from databases and from web sites. We call the result SEMCAT. It currently recognizes seventy-five categories and contains about five million name strings distributed over those categories. We have experimented with probabilistic context free grammars and Markov models of text strings in an attempt to learn how to recognize the entities in different categories. In order to improve performance we have developed a new model term a Priority Model for name recognition. This model allows us to categorize names as gene/protein names with an F-score of 0.96 and better then what we were able to achieve with either a language model of a probabilistic context free grammar. We are currently using this to create features for using in a conditional random fields approach to gene/protein name recognition and are achieving about an 0.83 F-score.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A DOCUMENT PROCESSING SYSTEM
-
批准号:6111059
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6111081
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6290494
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical phrase extraction techniques in natural language databases.
-
批准号:6228047
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6546806
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme recognition in document collections.
-
批准号:6432765
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:6843586
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:7316268
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7148024
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7148023
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:7594456
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6843558
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:6843556
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
RETRIEVAL TASK--COMPARE GROUP/INDIVIDUAL PERFORMANCE--SUBJECT EXPERTS/UNTRAINED
-
批准号:6432759
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
A Document Processing System
-
批准号:6546807
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
AUTOMATIC BAYESIAN METHODS IN TEXT RETRIEVAL
-
批准号:6432747
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7316262
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Statistical Phrase Extraction Techniques In Natural Lang
-
批准号:7316265
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections
-
批准号:7148039
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
Theme Recognition In Document Collections.
-
批准号:7594468
-
项目类别:
-
资助金额:$5.3万
-
财政年份:--
-
负责人:willy john wilbur
-
依托单位:
海外基金