Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
批准号:
8558117
负责人:
Willy Wilbur
金额:
$26.03万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AccountingAddressBooksCellsCharacteristicsCollectionDataElectronicsEvaluationFrequenciesGeneric DrugsGoalsLifeLiteratureMachine LearningMethodsMutationOutcomeProcessPropertyPubMedRecordsRetrievalSchemeSpeechStagingTechniquesTextTextbooksTimeTrainingUMLS Metathesaurusbasedesignimprovedindexinginsightphrasesresponsestem
中文摘要
点击翻译按钮获取中文摘要
英文摘要
1) Electronic Textbook and PubMed Central Indexing
Current processing of the electronic textbook material involves a number of steps designed to produce the most meaningful phrases in the text to be used as reference points. The first task is to identify grammatically reasonable phrases. We use a version of the Brill transformation based tagger, rewritten in C++, for part-of- speech tagging. This forms the basis for determining grammatically reasonable phrases. There is a significant post processing step that removes phrases that involve inappropriate references to context (e.g., different cells, final mutation). After finding grammatically reasonable phrases we attempt to eliminate those that are too common or generic to be useful (e.g., significant result, short time). The next step is to compare a phrase with previously rated phrases that have been collected over the life of the project. The final stage is to estimate the importance of a phrase in the passage where it is found in a textbook. Such an estimate is based on the frequency of the phrase and the size of the passage compared with the frequency of the phrase throughout the book and the overall size of the book. In order to improve such an estimate we attempt to take account of the phrase or any phrase that represents the same concept. For this purpose we use the UMLS Metathesaurus and also stemming and combine these two approaches into a consistent picture of the concept as it occurs in the text. The result of this processing is a scored list of phrase-book section pairs for each textbook. These are used to guide the response of general searching in the books. When a user types in a phrase that is on our curated list the first results given are the highly rated book sections for that phrase. We are now applying a similar indexing scheme to the text of articles in PMCentral. This allows us to give a list of highly rated phrases for each article as an enhanced reference point for searchers.
2) A significant fraction of queries in PubMed are multiterm queries and PubMed generally handles them as a Boolean conjunction of the terms. However, analysis of queries in PubMed indicates that many such queries are meaningful phrases, rather than simply collections of terms. We have examined whether or not it makes a difference, in terms of retrieval quality, if such queries are interpreted as a phrase or as a conjunction of query terms. And, if it does, what is the optimal way of searching with such queries. To address the question, we developed an automated retrieval evaluation method, based on machine learning techniques, that enables us to evaluate and compare various retrieval outcomes. We show that classes of records that contain all the search terms, but not the phrase, qualitatively differ from the class of records containing the phrase. We also show that the difference is systematic, depending on the proximity of query terms to each other within the record. Based on these results, one can establish the best retrieval order for the records. Our findings are consistent with studies in proximity searching. The important insight here for indexing is that in some cases where the words of a phrase occur in text, but not as the phrase, the phrase may still be an appropriate concept to use in indexing the text.
3) Currently we are studying how good phrases can be recognized by their characteristics, such as frequency, tendency to be repeated in documents where they occur, and other numerical properties. These features allow one to predict which phrases are of high quality. We have found such predictions to be useful in studying different kinds of terms that may appear in text and how an ontoloogy might be extracted from text.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Document Processing System
-
批准号:8344939
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:8344960
-
项目类别:
-
资助金额:$23.98万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8558105
-
项目类别:
-
资助金额:$56.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Natural Language Processing Techniques To Enhance Information Access.
-
批准号:8943224
-
项目类别:
-
资助金额:$56.15万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7969244
-
项目类别:
-
资助金额:$77.4万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8149591
-
项目类别:
-
资助金额:$13.71万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8149592
-
项目类别:
-
资助金额:$17.63万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8149602
-
项目类别:
-
资助金额:$47.01万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:9160906
-
项目类别:
-
资助金额:$43.97万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:7969199
-
项目类别:
-
资助金额:$20.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
-
批准号:8344948
-
项目类别:
-
资助金额:$59.96万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:8344950
-
项目类别:
-
资助金额:$17.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
A Document Processing System
-
批准号:8943215
-
项目类别:
-
资助金额:$18.72万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:7969197
-
项目类别:
-
资助金额:$12.9万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Bayesian Methods In Text Retrieval
-
批准号:8344938
-
项目类别:
-
资助金额:$7.99万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:9160916
-
项目类别:
-
资助金额:$14.32万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
-
批准号:9160928
-
项目类别:
-
资助金额:$12.27万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
-
批准号:7735088
-
项目类别:
-
资助金额:$29.31万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
PubMed Query Log Analysis and Use in Access Enhancement
-
批准号:8177730
-
项目类别:
-
资助金额:$48.97万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
Free Text Gene Name Recognition
-
批准号:8149604
-
项目类别:
-
资助金额:$19.59万
-
财政年份:--
-
负责人:Willy Wilbur
-
依托单位:
海外基金