Beyond Abstracts: Issues in Mining Full Texts
Beyond Abstracts: Issues in Mining Full Texts
批准号:
7287359
负责人:
LAWRENCE E HUNTER
金额:
$35.06万
依托单位国家:
美国
项目类别:
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-09-15 至 2009-09-14
关键词:
AdoptionAffectAgreementArtsBiomedical ResearchBody of uterusCollectionComputational TechniqueDataDevelopmentEvaluationGoldGrowthHumanJudgmentLanguageLeadLiteratureMachine LearningMemoryMethodsMetricMiningModelingMolecular BiologyNumbersPatternPeer ReviewPerformancePlayProcessPublic HealthPublishingRangeRepresentations, Knowledge (Computer)ResearchRetrievalReview LiteratureRoleSamplingSchemeScientistStandards of Weights and MeasuresSystemTechniquesTechnologyTestingTextTrainingWorkabstractingbaseconceptimprovedinformation organizationinterestjournal articlelanguage processingnovel strategiesprototypesizetechnology developmenttext searchingtool
中文摘要
描述(由申请人提供):
英文摘要
DESCRIPTION (provided by applicant):
Biomedical language processing, the application of computational techniques to human-generated texts in biomedicine, is an increasingly important enabling technology for basic and applied biomedical research. The exponential growth of the peer-reviewed literature and the breakdown of disciplinary boundaries associated with high-throughput techniques have increased the importance of automated tools for keeping scientists abreast of all of the published material relevant to their work. However, despite decades of research, the performance of state-of-the-art tools for basic language processing tasks like information extraction and document retrieval remain below the level necessary for adequate utility and widespread adoption of this technology. The development, performance and evaluation of text mining systems depend crucially on the availability of appropriate corpora: collections of representative documents that have been annotated with human judgments relevant to a language-processing task. Corpora play two roles in the development of this technology: first, they act as "gold standards" by which alternative automated methods can be fairly compared, and second, they provide data for the training of statistical and machine learning systems that create empirical models of patterns in language use. The conventional view is that corpora are neutral, random samples of the domain of interest. Our preliminary work suggests that the restrictions in size, quality, genre, and representational schema of the small number of existing corpora are themselves a critical limiting factor for near-term breakthroughs in biomedical text processing technology. Therefore, we propose to test the following hypothesis: Creation of large, high-quality, biomedical corpora from multiple genres will lead to significant improvements in the performance of biomedical text mining systems and the creation of new approaches to text mining tasks. Specific aims include constructing several large corpora covering a range of genres and incorporating a rich knowledge representation; identifying factors that affect differential performance on full text versus abstracts; and developing new methods for language processing, especially of full text. Because improvements in the ability to automatically extract information from many textual genres will assist scientists and clinicians in the crucial task of keeping up with the burgeoning biomedical literature, the potential public health impact is quite large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
High Performance Text Mining for Translator
-
批准号:10334356
-
项目类别:
-
资助金额:$47.12万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Scientific Questions: A New Target for Biomedical NLP
-
批准号:10223438
-
项目类别:
-
资助金额:$45.31万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Scientific Questions: A New Target for Biomedical NLP
-
批准号:10454968
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
High Performance Text Mining for Translator
-
批准号:10548337
-
项目类别:
-
资助金额:$46.61万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9526127
-
项目类别:
-
资助金额:$9.98万
-
财政年份:2017
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Automated Literature Mining for Validation of High-Throughput Function Prediction
-
批准号:7843633
-
项目类别:
-
资助金额:$71.14万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Construction of a Full Text Corpus for Biomedical Text Mining
-
批准号:7872692
-
项目类别:
-
资助金额:$6.6万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7824978
-
项目类别:
-
资助金额:$44.56万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7877947
-
项目类别:
-
资助金额:$47.07万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:8261523
-
项目类别:
-
资助金额:$87.05万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7364235
-
项目类别:
-
资助金额:$63.16万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7502636
-
项目类别:
-
资助金额:$64.09万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7684604
-
项目类别:
-
资助金额:$63.91万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9264193
-
项目类别:
-
资助金额:$49.42万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:8681518
-
项目类别:
-
资助金额:$76.45万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:8133187
-
项目类别:
-
资助金额:$21.52万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9105400
-
项目类别:
-
资助金额:$79.54万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Construction of a Full Text Corpus for Biomedical Text Mining
-
批准号:7301251
-
项目类别:
-
资助金额:$13.04万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7928868
-
项目类别:
-
资助金额:$60.5万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7457685
-
项目类别:
-
资助金额:$49.78万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
海外基金