Construction of a Full Text Corpus for Biomedical Text Mining
Construction of a Full Text Corpus for Biomedical Text Mining
批准号:
7673720
负责人:
LAWRENCE E HUNTER
金额:
$14.29万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-15 至 2010-09-14
关键词:
AddressAgreementBiologicalBiologyBody of uterusCollectionCommunitiesConsensusDataDatabasesDevelopmentElementsEnsureFeedbackGenesGoldGrowthHumanLightLinguisticsLiteratureMEDLINEMachine LearningManualsMeasuresMetricMonitorNatural Language ProcessingNatureOntologyOutputProblem SolvingProceduresProcessPublicationsPublished CommentResearchResearch PersonnelResourcesRoleSchemeSeriesStatistical ModelsStructureSystemTestingTextTrainingUnified Medical Language SystemWorkabstractingbasedesignexperienceindexinginformation organizationinnovationjournal articlelanguage processingmodel organisms databasesprogramsquality assurancetext searchingtrend
中文摘要
描述(由申请人提供):
有一个被证明的社区需要一个注释的语料库组成的生物医学期刊文章的全文。有很多理由相信,阻碍当今生物医学语言处理进步的限速因素是缺乏正确的专业注释数据。带注释的语料库是文本的集合,其中包含与特定文本元素相关的含义或结构的信息。注释语料库在两个方面是生物医学自然语言处理研究的重要组成部分。首先,大多数当代语言处理方法至少部分依赖于机器学习或统计模型。这样的系统必须在具有已知输出的示例集上进行“训练”,因此注释语料库提供了对现代NLP系统的构建至关重要的训练数据。第二,注释语料库提供了黄金标准,通过它来评估特定文本挖掘任务的各种方法。由于它们在训练和测试语言处理系统中的核心作用,注释语料库的设计和操作创建的质量对这些系统可以完成的工作产生了根本性的限制。尽管在注释摘要方面已经做了有价值的工作,但从文本挖掘的角度来看,摘要和全文文章之间存在重要差异,全文期刊文章的注释可以忽略不计。生物(特别是模式生物数据库管理)社区和文本挖掘社区的工作人员都独立地指出,如果生物医学世界能够充分利用文本挖掘,处理科学出版物的全文的重要性。我们建议建立一个大型的,充分注释的语料库,包括生物医学期刊文章的全文。此外,以前的生物医学语料库注释工作经常利用专门的本体,这限制了它们在创建它们的组之外的效用。我们将确保社区的可接受性注释方面的社区共识的本体论,如基因本体论和UMLS。由于这项任务涉及昂贵的人力,效率是创建语料库的关键问题。出于这个原因,我们建议建立一个团队,其中包括迄今为止最大的语义注释语料库的构建者,模式生物数据库的先驱之一,以及经验丰富的语言学和领域专家注释者的骨干。
英文摘要
DESCRIPTION (provided by applicant):
There is a demonstrated community need for an annotated corpus consisting of the full texts of biomedical journal articles. There are many reasons to believe that the rate-limiting factor impeding progress in biomedical language processing today is the lack of availability of the right kind of expertly annotated data. An annotated corpus is a collection of texts with information about the meaning or structure associated with particular textual elements. Annotated corpora are a critical component of biomedical natural language processing research in two ways. First, most contemporary approaches to language processing rely at least in part on machine learning or statistical models. Such systems must be "trained" on sets of examples with known outputs, so annotated corpora provide the training data vital to the construction of modern NLP systems. Second, annotated corpora provide the gold standard by which various approaches to particular text mining tasks are evaluated. Due to their central roles in training and testing language processing systems, the quality of the design and operational creation of annotated corpora place fundamental limits on what can be accomplished with such systems. Although there has been valuable work done on annotating abstracts, there are important differences between abstracts and full-text articles from a text mining perspective, and annotation of full-text journal articles has been negligible. Workers in both the biological (especially model organism database curation) community and the text mining community have independently pointed out the importance of processing the full text of scientific publications if the biomedical world is to be able to fully utilize text mining. We propose to build a large, fully annotated corpus consisting of full texts of biomedical journal articles. Additionally, previous biomedical corpus annotation efforts have often utilized ad hoc ontologies that have limited their utility outside of the groups that created them. We will ensure community acceptability by annotating with respect to community-consensus ontologies such as the Gene Ontology and the UMLS. Since the task involves expensive human labor, efficiency is a key issue in creating corpora. For this reason, we propose to build a team that includes the builder of the largest semantically annotated corpus to date, one of the pioneers of the model organism databases, and an already-assembled cadre of experienced linguistic and domain-expert annotators.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
High Performance Text Mining for Translator
-
批准号:10334356
-
项目类别:
-
资助金额:$47.12万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Scientific Questions: A New Target for Biomedical NLP
-
批准号:10223438
-
项目类别:
-
资助金额:$45.31万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Scientific Questions: A New Target for Biomedical NLP
-
批准号:10454968
-
项目类别:
-
资助金额:$44.52万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
High Performance Text Mining for Translator
-
批准号:10548337
-
项目类别:
-
资助金额:$46.61万
-
财政年份:2020
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9526127
-
项目类别:
-
资助金额:$9.98万
-
财政年份:2017
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Automated Literature Mining for Validation of High-Throughput Function Prediction
-
批准号:7843633
-
项目类别:
-
资助金额:$71.14万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Construction of a Full Text Corpus for Biomedical Text Mining
-
批准号:7872692
-
项目类别:
-
资助金额:$6.6万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7824978
-
项目类别:
-
资助金额:$44.56万
-
财政年份:2009
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:8261523
-
项目类别:
-
资助金额:$87.05万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7364235
-
项目类别:
-
资助金额:$63.16万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7877947
-
项目类别:
-
资助金额:$47.07万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7502636
-
项目类别:
-
资助金额:$64.09万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7684604
-
项目类别:
-
资助金额:$63.91万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:8681518
-
项目类别:
-
资助金额:$76.45万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9264193
-
项目类别:
-
资助金额:$49.42万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:8133187
-
项目类别:
-
资助金额:$21.52万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Colorado Biomedical Informatics Training Program
-
批准号:9105400
-
项目类别:
-
资助金额:$79.54万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Construction of a Full Text Corpus for Biomedical Text Mining
-
批准号:7301251
-
项目类别:
-
资助金额:$13.04万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Ontologies and Biomedical Language Processing
-
批准号:7928868
-
项目类别:
-
资助金额:$60.5万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
Computational Bioscience Program Training Grant
-
批准号:7457685
-
项目类别:
-
资助金额:$49.78万
-
财政年份:2007
-
负责人:LAWRENCE E HUNTER
-
依托单位:
海外基金