Incorporating Image-based Features into Biomedical Document Classification
Incorporating Image-based Features into Biomedical Document Classification
批准号:
9762175
负责人:
Georgeta-Elisabeta Marai
金额:
$46.3万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-14 至 2021-08-31
关键词:
AddressAreaBiological PhenomenaCategoriesClassificationCollaborationsComputer-Assisted Image AnalysisCuesDataData SetDatabasesDevelopmentDiseaseDrug TargetingFailureFluorescence MicroscopyFoundationsGelGene ExpressionGene MutationGene ProteinsGenomicsGeometryGoalsGrainHarvestImageImage AnalysisIndividualInformaticsInformation ResourcesInstitutesInvestigationLettersLiteratureMedicalMethodsMiningModelingMusMutationOutcomes ResearchPaperPhenotypePhysiciansPositioning AttributeProcessProteinsProteomicsPubMedPublicationsPublishingResearchResource InformaticsRetrievalRoleScanningSchemeScientistSecureShapesSolidSourceSpeedStructural ProteinSystemTextTextureTrainingWorkbasebioimagingbiomedical scientistdecision researchevaluation/testingexperienceexperimental studyimage processingimprovedindexingmouse genomemultimodalitynew therapeutic targetnovelprotein protein interactionprotein structuretext searchingtool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The proposed research aims to develop and advance tools for using image-data appearing in scientific publications, in addition to text, in order to support beneficial, targeted access to the biomedical literature. The number of biomedical publications grows at a rate of over one million new publications per year. Identifying relevant information requires scientists and physicians to scan daily through a myriad of papers. For scientific database curators (bio-curators, in organizations such as Jackson Labs or UniProt), the task is particularly onerous, as they must identify articles most significant to the database, locate within them high-quality evidence concerning disease, genes/proteins and mutations, and curate the findings in database entries along with references to relevant evidence in the articles. Notably, much of the evidence within publications lies in figures. Thus, images are rich and essential indicators for relevance.
While biomedical text mining tools are being developed to expedite search for information within publications, several competitive shared tasks underscored the need for more effective tools to overcome the bottleneck for bio-curation and for scientific discovery. Moreover, bio-curators point-out the importance of images as a key information source. While image analysis is an active research field, most current work on biomedical image processing focuses on image identification, understanding and indexing; Not on images as aids to document analysis. Similarly, most work on biomedical literature mining focuses on text alone. Thus, little has been done so far to utilize, in addition to text, images within publications that provide important cues about the relevance of the information embedded in articles.
Our premise, supported by bio-curators experience, is that information derived from images can (and should) be directly incorporated into biomedical document retrieval and classification, and will improve accurate identification of relevant articles (for a given user’s needs) while pin-pointing significant evidence within them. We will comprehensively identify, develop and compare informative image-features, develop methods and tools for representing both images and documents based on such features, and introduce means to effectively integrate image-based data into the text-based document classification process. The work will comprise the following fundamental tasks: A) Building robust tools for harvesting images from PDF articles and segmenting compound figures into individual image-panels; B) Identification and investigation of highly-informative features for biomedical image-representation, and categorization of biomedical images into significant types and classes; C) Effective representation of documents using text and image, and integration of text-based and image-based classifiers. We anchor our research in genuine needs, secure access to much image data, and strive for broad-applicability of the results, by working within several broad and diverse curation-areas within institutes with which we collaborate: Evidence for gene-expression & phenotypes in Mouse (Jackson Labs) and in worm (WormBase), and experimental evidence for protein-protein interaction (Protein Information Resource). The work on this project will result in new methods and tools that take advantage of both image- and text-data, facilitating more effective and focused retrieval and mining, thus better supporting bio-curation and data-intensive biomedical discovery.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Incorporating Image-based Features into Biomedical Document Classification
-
批准号:9457095
-
项目类别:
-
资助金额:$48.82万
-
财政年份:2017
-
负责人:Georgeta-Elisabeta Marai
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: